PySpark写入数据字典时触发TypeError的技术求助
解决Spark DataFrame统计结果写入文件的TypeError问题
问题原因
你的process函数仅打印统计结果但没有返回任何值(Python函数默认返回None),而file.write()要求传入字符串类型参数,因此触发TypeError: write() argument must be str, not None。
解决方案
需要修改process函数,让它收集所有列的统计数据并返回可写入文件的字符串格式,同时可以优化代码减少冗余、提升效率:
- 移除不必要的
func类,直接遍历DataFrame的列名列表 - 提前计算总记录数,避免重复调用
df.count()(每次调用都会触发Spark作业,影响效率) - 让
process函数返回整合后的统计结果字符串(推荐用JSON格式,可读性更强) - 使用
with语句管理文件,自动处理文件关闭,更安全
修改后的完整代码
from pyspark.sql import SparkSession import pyspark.sql.functions as F import json # 初始化SparkSession spark = SparkSession.builder.appName('SparkByExamples.com').getOrCreate() data = [(1,"Aj",None,"ABC"), (2,"Kishore","N","DEF"), (3,"Kishore","P",None),(4,"Naveen","N","XYZ")] columns = ["empid","empfname","emplname","empjobcode"] df = spark.createDataFrame(data=data, schema=columns) df.show() def process(column_list): total_count = df.count() # 提前计算总记录数,避免重复计算 stats_result = {} for col_name in column_list: # 计算各统计指标 null_count = df.where(F.col(col_name).isNull()).count() not_null_count = total_count - null_count distinct_count = df.select(col_name).distinct().count() # 填充统计字典 stats_result[col_name] = { "Null count": null_count, "Percentage of Null": round((null_count / total_count) * 100, 2), "Not-null count": not_null_count, "Percentage of Not-null": round((not_null_count / total_count) * 100, 2), "Distinct value count": distinct_count, "Percentage of Distinct value": round((distinct_count / total_count) * 100, 2) } # 将统计结果转为格式化的JSON字符串 return json.dumps(stats_result, indent=4) # 调用函数并写入文件 with open('New.txt', 'w') as file: file.write(process(df.columns))
代码说明
- 总记录数预计算:
total_count = df.count()只执行一次,后续所有占比计算都复用这个值,减少Spark作业次数 - 统计结果整合:用字典
stats_result统一存储所有列的统计数据,结构更清晰 - JSON格式化:
json.dumps(..., indent=4)将字典转为带缩进的JSON字符串,写入文件后可读性强 - with语句:自动处理文件打开/关闭,避免资源泄漏
写入文件后的示例内容
{ "empid": { "Null count": 0, "Percentage of Null": 0.0, "Not-null count": 4, "Percentage of Not-null": 100.0, "Distinct value count": 4, "Percentage of Distinct value": 100.0 }, "empfname": { "Null count": 0, "Percentage of Null": 0.0, "Not-null count": 4, "Percentage of Not-null": 100.0, "Distinct value count": 3, "Percentage of Distinct value": 75.0 }, "emplname": { "Null count": 1, "Percentage of Null": 25.0, "Not-null count": 3, "Percentage of Not-null": 75.0, "Distinct value count": 3, "Percentage of Distinct value": 75.0 }, "empjobcode": { "Null count": 1, "Percentage of Null": 25.0, "Not-null count": 3, "Percentage of Not-null": 75.0, "Distinct value count": 4, "Percentage of Distinct value": 100.0 } }
内容的提问来源于stack exchange,提问作者AAA6
相关产品推荐
相关产品推荐

