Synapse代码无法删除DataFrame中replaceCols数组指定列的问题
问题描述
给定列名列表:
replaceCols=[ "LEI", "Entity_LegalName", "Entity_LegalAddress_FirstAddressLine", "Entity_LegalAddress_City", "Entity_LegalAddress_Country", "Entity_HeadquartersAddress_FirstAddressLine", "Entity_HeadquartersAddress_City", "Entity_HeadquartersAddress_Country", "Entity_RegistrationAuthority_RegistrationAuthorityID", "Entity_LegalJurisdiction", "Entity_LegalForm_EntityLegalFormCode", "Entity_EntityStatus", "Registration_InitialRegistrationDate", "Registration_LastUpdateDate", "Registration_RegistrationStatus", "Registration_NextRenewalDate", "Registration_ManagingLOU", "Registration_ValidationSources" ]
读取CSV文件得到的DataFrame df 包含上述列及其他更多列,执行以下代码:
df_W = df.drop(str(col(c).alias(c) for c in replaceCols))
代码语法无报错,但无法删除replaceCols中指定的列。
问题分析与解决
这段代码失效的核心原因是参数传递错误:
str(col(c).alias(c) for c in replaceCols)会把生成器对象直接转为字符串,结果是类似<generator object <genexpr> at 0x...>的无效内容,Spark无法识别这是要删除的列。
有两种正确的写法:
- 直接传入列名列表(最简洁):
df_W = df.drop(*replaceCols)
这里的*是解包操作符,将列表中的每个列名作为独立参数传递给drop方法。
- 传入列对象列表(适用于需要对列做额外处理的场景):
from pyspark.sql.functions import col df_W = df.drop(*[col(c) for c in replaceCols])
内容的提问来源于stack exchange,提问作者Wasim Syed
相关产品推荐
相关产品推荐

