You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在PySpark DataFrame中选择列并转换类型,同时保留ID字符串列

你可以直接在select方法的列列表中拼接ID列,不需要修改原有批量转换的逻辑,写法如下:

from pyspark.sql.functions import col
from pyspark.sql.types import DoubleType

df_num2 = df_num1.select(
    # 把ID列放在最前面,也可以根据需求调整拼接顺序放到最后
    [col("ID")] + [col(c).cast(DoubleType()) for c in num_columns]
)

如果你的num_columns列表中可能包含ID列,为了避免重复选择同名列报错,可以先做一次过滤:

# 先过滤掉num_columns中的ID列
num_columns_filtered = [c for c in num_columns if c != "ID"]
df_num2 = df_num1.select(
    [col("ID")] + [col(c).cast(DoubleType()) for c in num_columns_filtered]
)

这种写法没有额外的计算开销,完全适配大体积PySpark DataFrame的处理场景。

内容的提问来源于stack exchange,提问作者Estelle D

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 12:24:02