Azure Databricks中Python实现DataFrame Interest列拆分展开
Azure Databricks中拆分Interest列并处理字段的实现方法
问题场景
原始DataFrame结构:
id Country Interest 00Q7 Russian Digestion;Destillation
期望处理后的结果:
id Country Int Interest 00Q7 Russ Digestion Digestion;Destillation 00Q7 Russ Destillation Digestion;Destillation
实现步骤
1. 导入PySpark函数模块
在Databricks Python Notebook中,先导入所需的PySpark函数:
from pyspark.sql import functions as F
2. 拆分并展开Interest列
使用split将Interest列按分号分割为数组,再用explode将数组元素拆分为单独行:
# 假设原始DataFrame名为df df_step1 = df.withColumn("Int", F.explode(F.split(F.col("Interest"), ";")))
3. 截断Country字段
提取Country列的前4个字符作为新的Country值:
df_step2 = df_step1.withColumn("Country", F.substring(F.col("Country"), 1, 4))
4. 调整列顺序(可选)
如果需要匹配目标结果的列顺序,重新排列列:
df_final = df_step2.select("id", "Country", "Int", "Interest")
完整可运行代码
from pyspark.sql import functions as F # 模拟原始数据(实际场景替换为读取数据源的逻辑,如spark.read.table等) sample_data = [("00Q7", "Russian", "Digestion;Destillation")] df = spark.createDataFrame(sample_data, schema=["id", "Country", "Interest"]) # 执行全流程处理 df_final = df.withColumn("Int", F.explode(F.split(F.col("Interest"), ";"))) \ .withColumn("Country", F.substring(F.col("Country"), 1, 4)) \ .select("id", "Country", "Int", "Interest") # 查看处理结果 df_final.show(truncate=False)
内容的提问来源于stack exchange,提问作者Rohit Kulkarni
相关产品推荐
相关产品推荐

