You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark写入数据字典时触发TypeError的技术求助

解决Spark DataFrame统计结果写入文件的TypeError问题

问题原因

你的process函数仅打印统计结果但没有返回任何值(Python函数默认返回None),而file.write()要求传入字符串类型参数,因此触发TypeError: write() argument must be str, not None。

解决方案

需要修改process函数,让它收集所有列的统计数据并返回可写入文件的字符串格式,同时可以优化代码减少冗余、提升效率:

  • 移除不必要的func类,直接遍历DataFrame的列名列表
  • 提前计算总记录数,避免重复调用df.count()(每次调用都会触发Spark作业,影响效率)
  • 让process函数返回整合后的统计结果字符串(推荐用JSON格式,可读性更强)
  • 使用with语句管理文件,自动处理文件关闭,更安全

修改后的完整代码

from pyspark.sql import SparkSession
import pyspark.sql.functions as F
import json

# 初始化SparkSession
spark = SparkSession.builder.appName('SparkByExamples.com').getOrCreate()
data = [(1,"Aj",None,"ABC"), (2,"Kishore","N","DEF"), (3,"Kishore","P",None),(4,"Naveen","N","XYZ")]
columns = ["empid","empfname","emplname","empjobcode"]
df = spark.createDataFrame(data=data, schema=columns)
df.show()

def process(column_list):
    total_count = df.count()  # 提前计算总记录数,避免重复计算
    stats_result = {}
    
    for col_name in column_list:
        # 计算各统计指标
        null_count = df.where(F.col(col_name).isNull()).count()
        not_null_count = total_count - null_count
        distinct_count = df.select(col_name).distinct().count()
        
        # 填充统计字典
        stats_result[col_name] = {
            "Null count": null_count,
            "Percentage of Null": round((null_count / total_count) * 100, 2),
            "Not-null count": not_null_count,
            "Percentage of Not-null": round((not_null_count / total_count) * 100, 2),
            "Distinct value count": distinct_count,
            "Percentage of Distinct value": round((distinct_count / total_count) * 100, 2)
        }
    
    # 将统计结果转为格式化的JSON字符串
    return json.dumps(stats_result, indent=4)

# 调用函数并写入文件
with open('New.txt', 'w') as file:
    file.write(process(df.columns))

代码说明

  • 总记录数预计算:total_count = df.count()只执行一次,后续所有占比计算都复用这个值,减少Spark作业次数
  • 统计结果整合:用字典stats_result统一存储所有列的统计数据,结构更清晰
  • JSON格式化:json.dumps(..., indent=4)将字典转为带缩进的JSON字符串,写入文件后可读性强
  • with语句:自动处理文件打开/关闭,避免资源泄漏

写入文件后的示例内容

{
    "empid": {
        "Null count": 0,
        "Percentage of Null": 0.0,
        "Not-null count": 4,
        "Percentage of Not-null": 100.0,
        "Distinct value count": 4,
        "Percentage of Distinct value": 100.0
    },
    "empfname": {
        "Null count": 0,
        "Percentage of Null": 0.0,
        "Not-null count": 4,
        "Percentage of Not-null": 100.0,
        "Distinct value count": 3,
        "Percentage of Distinct value": 75.0
    },
    "emplname": {
        "Null count": 1,
        "Percentage of Null": 25.0,
        "Not-null count": 3,
        "Percentage of Not-null": 75.0,
        "Distinct value count": 3,
        "Percentage of Distinct value": 75.0
    },
    "empjobcode": {
        "Null count": 1,
        "Percentage of Null": 25.0,
        "Not-null count": 3,
        "Percentage of Not-null": 75.0,
        "Distinct value count": 4,
        "Percentage of Distinct value": 100.0
    }
}

内容的提问来源于stack exchange,提问作者AAA6

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 07:31:19