Python自定义myfunccsvtojson实现CSV转表头为键的行分隔JSON
CSV转行分隔JSON函数实现
需求规则
- 编写名为
myfunccsvtojson的函数,接收CSV文件路径作为入参,生成支持流式读取的行分隔JSON文件 - 输出文件名基于输入CSV文件名自动生成,禁止硬编码:输入
Myfilename.csv输出Myfilename.json,输入File2.csv输出File2.json - 输入CSV共10000行(含表头行),输出JSON文件需对应生成9999行有效数据
- 转换规则:
- CSV表头作为每条JSON的键,对应行字段为JSON值
- CSV中
\N空值统一转为JSON的null - 年份类数值字段转为数字类型
样例参考
输入CSV样例
nconst,primaryName,birthYear,deathYear,primaryProfession,knownForTitles nm0000001,Fred Astaire,1899,1987,"soundtrack,actor,miscellaneous","tt0072308,tt0043044,tt0050419,tt0053137" nm0000002,Lauren Bacall,1924,2014,"actress,soundtrack","tt0071877,tt0038355,tt0117057,tt0037382" nm0000003,Brigitte Bardot,1934,\N,"actress,soundtrack,producer","tt0057345,tt0059956,tt0049189,tt0054452"
输出JSON行样例
{"nconst":"nm0000001","primaryName":"Fred Astaire","birthYear":1899,"deathYear":1987,"primaryProfession":"soundtrack,actor,miscellaneous","knownForTitles":"tt0072308,tt0043044,tt0050419,tt0053137"} {"nconst":"nm0000002","primaryName":"Lauren Bacall","birthYear":1924,"deathYear":2014,"primaryProfession":"actress,soundtrack","knownForTitles":"tt0071877,tt0038355,tt0117057,tt0037382"} {"nconst":"nm0000003","primaryName":"Brigitte Bardot","birthYear":1934,"deathYear":null,"primaryProfession":"actress,soundtrack,producer","knownForTitles":"tt0057345,tt0059956,tt0049189,tt0054452"}
原有代码错误点
- 同一文件句柄同时初始化
csv.DictReader和csv.reader,两个读取器共用文件指针,会导致读取内容错位、表头获取异常 - 循环语法不符合Python规则:
for i,Value in range(len(keys)), csv_data无法正确并行遍历键列表和行数据,属于逻辑语法错误 - 缺失要求的类型转换逻辑:未处理
\N空值转null的规则,也未对年份字段做数字类型转换 - 重复实现表头映射逻辑:
csv.DictReader本身内置了表头到行字段的映射能力,不需要手动读取表头再循环赋值
完整实现代码
import csv import json import os def myfunccsvtojson(csv_path): # 自动生成输出JSON路径,无硬编码,适配任意输入文件名 file_root, _ = os.path.splitext(csv_path) json_output_path = f"{file_root}.json" # 流式读写,逐行处理,不加载全量数据到内存 with open(csv_path, "r", encoding="utf-8") as csv_f, \ open(json_output_path, "w", encoding="utf-8") as json_f: # DictReader自动将第一行表头作为每行字典的键,无需手动映射 reader = csv.DictReader(csv_f) for row in reader: processed_row = {} for key, value in row.items(): # \N空值转None,JSON序列化后自动为null if value == r"\N": processed_row[key] = None # 字段名含Year的年份字段转整数类型 elif "Year" in key and value.isdigit(): processed_row[key] = int(value) # 其余字段保留原字符串值,带引号的逗号内容会被csv模块自动正确解析 else: processed_row[key] = value # 逐行写入JSON,符合行分隔格式要求 json_f.write(json.dumps(processed_row, ensure_ascii=False) + "\n")
逻辑说明
- 文件名生成通过
os.path.splitext拆分输入路径的文件名和后缀,直接替换后缀为.json,支持任意合法文件路径输入 - 全程采用逐行读取、逐行转换、逐行写入的流式处理逻辑,内存占用恒定,不受CSV文件大小影响,输出的行分隔JSON文件天然支持流式读取
DictReader自动跳过表头行,仅遍历处理数据行,10000行输入(含表头)会恰好输出9999行JSON数据- 值转换逻辑完全匹配规则:带双引号的逗号分隔字段(如职业列表、作品ID列表)会被csv模块自动识别为单个字段,不会错误拆分
内容的提问来源于stack exchange,提问作者sumanth shetty
相关产品推荐
相关产品推荐

