如何将从DataFrame获取的字符串转换为另一个DataFrame?
问题描述
我在Parquet文件中存储了类似如下的列表数据:
mapping_in_parquet = [('filial','filial','S','string'),('doc','numero_do_documento','S','string'),('serie','serie_do_documento','S','string')]
通过以下代码将其提取到变量:
mapping = (df.select('mapping').distinct().collect()[0][0])
但尝试用该变量创建DataFrame时出现错误:
from pyspark.sql.types import StructType, StructField, StringType schema = StructType([ StructField("fieldName", StringType(), True), StructField("alias", StringType(), True), StructField("column_active", StringType(), True), StructField("typeField", StringType(), True)]) df = (spark.createDataFrame(mapping, schema)) print(mapping)
报错信息:
StructType can not accept object '[' in type
注:直接在控制台使用硬编码的列表运行代码无报错,仅从Parquet列提取值到变量时出现该问题。
解决方案
问题根源是从Parquet提取的mapping并非Python列表,而是字符串形式的列表(例如内容为"[('filial','filial','S','string'),('doc','numero_do_documento','S','string'),('serie','serie_do_documento','S','string')]"),而spark.createDataFrame无法直接解析字符串格式的列表。
解决方法是使用Python的ast.literal_eval()工具将字符串转换为真实的Python列表:
import ast from pyspark.sql.types import StructType, StructField, StringType # 将字符串格式的列表转换为Python列表 mapping_list = ast.literal_eval(mapping) # 基于转换后的列表创建DataFrame schema = StructType([ StructField("fieldName", StringType(), True), StructField("alias", StringType(), True), StructField("column_active", StringType(), True), StructField("typeField", StringType(), True)]) df = spark.createDataFrame(mapping_list, schema)
验证方式
转换后可以通过type(mapping_list)检查变量类型,确认其为<class 'list'>而非字符串,确保后续创建DataFrame的操作正常执行。
内容的提问来源于stack exchange,提问作者SrKartcheski
相关产品推荐
相关产品推荐

