如何将PySpark DataFrame所有字段值转为大写且保留原有列名
问题原因
你写的列表推导仅生成了列转换逻辑的Column对象列表,没有传入DataFrame的select方法执行转换,因此得到的不是转换后的DataFrame。
正确实现代码
你当前数据集所有列均为字符串类型,直接传入转换列表到select方法即可:
from pyspark.sql import functions as F lastvalue_month = lastvalue_month.select([F.upper(F.col(col_name)).alias(col_name) for col_name in lastvalue_month.columns])
兼容非字符串列的优化方案
如果后续数据集可能加入数字、日期等非字符串类型列,直接对非字符串列执行upper会报错,可以增加类型判断只处理字符串列:
from pyspark.sql import functions as F from pyspark.sql.types import StringType # 筛选字符串类型列和其他类型列 str_cols = [field.name for field in lastvalue_month.schema.fields if isinstance(field.dataType, StringType)] other_cols = [col for col in lastvalue_month.columns if col not in str_cols] # 字符串列转大写,其他列保持原值 lastvalue_month = lastvalue_month.select( *other_cols, *[F.upper(F.col(col_name)).alias(col_name) for col_name in str_cols] )
内容的提问来源于stack exchange,提问作者Nabih Bawazir
相关产品推荐
相关产品推荐

