You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用同国家非空值填充PySpark DataFrame中Continent列的空值?

PySpark 填充同一国家缺失的 Continent 字段

你的需求是用同一国家内已有的非空 Continent 值填充该国家内的空值,这里提供两种实用方法:

方法一:窗口函数(推荐)

利用窗口函数按国家分组,直接在原数据中提取非空值填充空记录,无需额外关联操作:

from pyspark.sql import Window
from pyspark.sql.functions import first, col

# 定义窗口:按Country分组,覆盖整个分组的行范围
window_spec = Window.partitionBy("Country").orderBy("Code").rowsBetween(Window.unboundedPreceding, Window.unboundedFollowing)

# 用第一个非空的Continent值填充同国家内的空值
filled_df = df.withColumn(
    "Continent",
    first(col("Continent"), ignoreNulls=True).over(window_spec)
)

# 查看结果
filled_df.show(truncate=False)

说明:

  • partitionBy("Country"):确保仅在同一国家的记录范围内处理
  • first(..., ignoreNulls=True):跳过空值,获取分组内第一个有效的Continent值,也可以用last()函数,效果一致
  • 窗口范围设置为整个分组,保证能获取到该国家所有记录中的非空值

方法二:分组映射+关联

先构建国家与大洲的映射表,再通过左关联填充空值:

from pyspark.sql.functions import col

# 生成国家-大洲映射表(过滤空值并去重)
continent_map = df.filter(col("Continent").isNotNull()).select("Country", "Continent").distinct()

# 左关联原表,用映射表的值填充空值
filled_df = df.join(continent_map, on="Country", how="left") \
    .withColumn("Continent", col("Continent_y")) \
    .drop("Continent_x", "Continent_y")

filled_df.show(truncate=False)

说明:

  • 映射表仅保留有有效大洲值的国家记录,去重后避免重复关联
  • 左关联保证原表所有记录都保留,没有有效大洲值的国家(如France)会继续保留Null

两种方法都能得到你预期的输出结果,窗口函数方法更高效,适合处理大数据量场景。

内容的提问来源于stack exchange,提问作者Mohammad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 17:51:30