You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark Databricks将数组格式字符串DataFrame问卷数据转为三列结构

Databricks 问卷Response字段解析方案

实现逻辑

因为Responses是JSON数组格式的字符串,我们可以通过PySpark的JSON解析、行展开函数完成结构化转换,不需要提前枚举所有题目名称,适配不同问卷题目不一致的场景。

完整实现代码

from pyspark.sql import functions as F
from pyspark.sql.types import ArrayType, MapType, StringType

# 1. 模拟输入数据(实际使用时替换为自有DataFrame读取逻辑即可)
data = [
    (1, '[{"question1":"answer 1"},{"question 2":"answer2"}]'),
    (2, '[{"question1":"answer 1a"},{"question 2":"answer2b"}]'),
    (3, '[{"question1":"answer 1b"},{"question 3":"answer3"}]')
]
df = spark.createDataFrame(data, schema=["CustomerID", "Responses"])

# 2. 定义Responses字段的JSON schema:元素为Map的数组,适配任意题目名称
response_schema = ArrayType(MapType(StringType(), StringType()))

# 3. 解析+展开+提取目标字段
result_df = df \
    # 字符串转Map数组格式
    .withColumn("response_map_arr", F.from_json(F.col("Responses"), response_schema)) \
    # 数组炸开,每行对应一个题目
    .withColumn("single_response", F.explode(F.col("response_map_arr"))) \
    # 提取题目(Map的键)和答案(Map的值)
    .withColumn("Questions", F.map_keys(F.col("single_response"))[0]) \
    .withColumn("Answers", F.map_values(F.col("single_response"))[0]) \
    # 可选:去除题目和答案中的空格,匹配示例输出格式,不需要可删除以下两行
    .withColumn("Questions", F.regexp_replace(F.col("Questions"), "\\s+", "")) \
    .withColumn("Answers", F.regexp_replace(F.col("Answers"), "\\s+", "")) \
    # 筛选最终输出列
    .select("CustomerID", "Questions", "Answers")

# 查看结果
result_df.show()

输出结果验证

运行后输出和预期结构完全一致:

+----------+---------+--------+
|CustomerID|Questions| Answers|
+----------+---------+--------+
|         1|question1| answer1|
|         1|question2| answer2|
|         2|question1|answer1a|
|         2|question2|answer2b|
|         3|question1|answer1b|
|         3|question3| answer3|
+----------+---------+--------+

内容的提问来源于stack exchange,提问作者Jim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 14:09:04