使用for循环定义PySpark数据结构变量时遇语法错误,寻求帮助
PySpark动态定义StructType语法错误解决
问题代码
colNames = ['colA', 'colB', 'colC', 'colD', 'colE'] tsfresh_feature_set = StructType( [ StructField('field1', StringType(), True), StructField('field2', StringType(), True), StructField(item, DoubleType(), False) for item in colNames ] )
报错信息
SyntaxError: invalid syntax File "<command-621368>", line 9 StructField(item, DoubleType(), False) for item in colNames ^ SyntaxError: invalid syntax
错误原因
Python列表中不能直接将普通元素与生成器表达式并列放置。原代码里的StructField(item, DoubleType(), False) for item in colNames是生成器表达式,无法作为单个元素插入列表,导致语法解析失败。
解决方案
方案1:列表拼接
将固定字段的列表与推导生成的字段列表拼接:
from pyspark.sql.types import StructType, StructField, StringType, DoubleType colNames = ['colA', 'colB', 'colC', 'colD', 'colE'] tsfresh_feature_set = StructType( [ StructField('field1', StringType(), True), StructField('field2', StringType(), True) ] + [StructField(item, DoubleType(), False) for item in colNames] )
方案2:解包列表推导式
用*操作符解包推导生成的列表,直接插入原列表中:
from pyspark.sql.types import StructType, StructField, StringType, DoubleType colNames = ['colA', 'colB', 'colC', 'colD', 'colE'] tsfresh_feature_set = StructType( [ StructField('field1', StringType(), True), StructField('field2', StringType(), True), *[StructField(item, DoubleType(), False) for item in colNames] ] )
两种方案都能动态生成包含指定特征列的StructType,解决语法错误问题。
内容的提问来源于stack exchange,提问作者Arun
相关产品推荐
相关产品推荐

