You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Databricks中Python实现DataFrame Interest列拆分展开

Azure Databricks中拆分Interest列并处理字段的实现方法

问题场景

原始DataFrame结构:

id      Country               Interest
00Q7    Russian               Digestion;Destillation

期望处理后的结果:

id    Country     Int                  Interest
00Q7  Russ        Digestion            Digestion;Destillation
00Q7  Russ        Destillation         Digestion;Destillation 

实现步骤

1. 导入PySpark函数模块

在Databricks Python Notebook中,先导入所需的PySpark函数:

from pyspark.sql import functions as F

2. 拆分并展开Interest列

使用split将Interest列按分号分割为数组,再用explode将数组元素拆分为单独行:

# 假设原始DataFrame名为df
df_step1 = df.withColumn("Int", F.explode(F.split(F.col("Interest"), ";")))

3. 截断Country字段

提取Country列的前4个字符作为新的Country值:

df_step2 = df_step1.withColumn("Country", F.substring(F.col("Country"), 1, 4))

4. 调整列顺序(可选)

如果需要匹配目标结果的列顺序,重新排列列:

df_final = df_step2.select("id", "Country", "Int", "Interest")

完整可运行代码

from pyspark.sql import functions as F

# 模拟原始数据(实际场景替换为读取数据源的逻辑,如spark.read.table等)
sample_data = [("00Q7", "Russian", "Digestion;Destillation")]
df = spark.createDataFrame(sample_data, schema=["id", "Country", "Interest"])

# 执行全流程处理
df_final = df.withColumn("Int", F.explode(F.split(F.col("Interest"), ";"))) \
             .withColumn("Country", F.substring(F.col("Country"), 1, 4)) \
             .select("id", "Country", "Int", "Interest")

# 查看处理结果
df_final.show(truncate=False)

内容的提问来源于stack exchange,提问作者Rohit Kulkarni

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 07:10:33