You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在CSV中设置主副标题行并生成对应PySpark DataFrame

处理带主副标题的CSV生成PySpark DataFrame

当CSV包含主标题和子标题两行时,直接用header=True只会将第一行设为列名,要实现主副标题组合的结构,可按以下步骤操作:

步骤1:读取CSV文件(不指定表头)

先将所有行都当作数据读取,后续再单独处理表头:

temp_df = spark.read.csv('ab.csv', header=False)

步骤2:提取并处理表头行

获取前两行(主标题行和子标题行),并处理主标题中可能存在的空值(对应Excel合并单元格导出后的情况):

# 获取前两行表头数据
header_rows = temp_df.limit(2).collect()
main_headers = header_rows[0]
sub_headers = header_rows[1]

# 填充主标题中的空值,让合并单元格对应的列继承同一主标题
def fill_empty_headers(headers):
    filled = []
    current_header = ""
    for h in headers:
        if h.strip() != "":
            current_header = h.strip()
        filled.append(current_header)
    return filled

filled_main_headers = fill_empty_headers(main_headers)

步骤3:生成组合列名

将主标题和子标题拼接成新的列名,格式如主标题_子标题:

new_columns = [f"{filled_main_headers[i]}_{sub_headers[i].strip()}" for i in range(len(filled_main_headers))]

步骤4:过滤数据行并设置新列名

移除前两行表头,将剩余数据行与新列名绑定:

from pyspark.sql.window import Window
from pyspark.sql.functions import row_number, lit

# 添加行号用于过滤表头行(虚拟排序仅为生成行号)
window = Window.orderBy(lit(1))
df_with_data = temp_df.withColumn("row_num", row_number().over(window)) \
                      .filter("row_num > 2") \
                      .drop("row_num")

# 应用新列名到数据行
final_df = df_with_data.toDF(*new_columns)

步骤5:(可选)转换数据类型

默认所有列都是字符串类型,可根据实际数据转换为对应类型:

from pyspark.sql.types import IntegerType, FloatType, StringType

# 示例:假设前两列为数值型,其余为字符串
type_mapping = [IntegerType(), FloatType()] + [StringType()]*(len(new_columns)-2)
final_df = final_df.select([final_df[col].cast(type_mapping[i]).alias(col) for i, col in enumerate(new_columns)])

验证结果

查看最终DataFrame的结构和数据:

final_df.printSchema()
final_df.show()

内容的提问来源于stack exchange,提问作者sumit agrawal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 17:39:24