我是 spark 新手。我試圖展平數據框,但未能透過「爆炸」做到這一點。
原始資料框架構如下:
id|approvaljson 1|[{"approvertype":"1st line manager","status":"approved"},{"approvertype":"2nd line manager","status":"approved"}] 2|[{"approvertype":"1st line manager","status":"approved"},{"approvertype":"2nd line manager","status":"rejected"}]
我需要將其轉換為以下架構?
id|approvaltype|status 1|1st line manager|approved 1|2nd line manager|approved 2|1st line manager|approved 2|2nd line manager|rejected
我已經嘗試過
df_exploded = df.withcolumn("approvaljson", explode("approvaljson"))
但是我得到了錯誤:
Cannot resolve "explode(ApprovalJSON)" due to data type mismatch: parameter 1 requires ("ARRAY" or "MAP") type, however, "ApprovalJSON" is of "STRING" type.;
首先將類似json 的字串解析為結構數組,然後使用inline
將數組分解為行和列
df1 = df.withcolumn("approvaljson", f.from_json("approvaljson", schema="array<struct<approvertype string, status string>>")) df1 = df1.select("id", f.inline('approvaljson'))
結果
df1.show() +---+----------------+--------+ | ID| ApproverType| Status| +---+----------------+--------+ | 1|1st Line Manager|Approved| | 1|2nd Line Manager|Approved| | 2|1st Line Manager|Approved| | 2|2nd Line Manager|Rejected| +---+----------------+--------+
以上是無法分解 Spark 資料框中的巢狀 JSON的詳細內容。更多資訊請關注PHP中文網其他相關文章!