You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python自定义myfunccsvtojson实现CSV转表头为键的行分隔JSON

CSV转行分隔JSON函数实现

需求规则

  • 编写名为myfunccsvtojson的函数,接收CSV文件路径作为入参,生成支持流式读取的行分隔JSON文件
  • 输出文件名基于输入CSV文件名自动生成,禁止硬编码:输入Myfilename.csv输出Myfilename.json,输入File2.csv输出File2.json
  • 输入CSV共10000行(含表头行),输出JSON文件需对应生成9999行有效数据
  • 转换规则:
    • CSV表头作为每条JSON的键,对应行字段为JSON值
    • CSV中\N空值统一转为JSON的null
    • 年份类数值字段转为数字类型

样例参考

输入CSV样例

nconst,primaryName,birthYear,deathYear,primaryProfession,knownForTitles
nm0000001,Fred Astaire,1899,1987,"soundtrack,actor,miscellaneous","tt0072308,tt0043044,tt0050419,tt0053137"
nm0000002,Lauren Bacall,1924,2014,"actress,soundtrack","tt0071877,tt0038355,tt0117057,tt0037382"
nm0000003,Brigitte Bardot,1934,\N,"actress,soundtrack,producer","tt0057345,tt0059956,tt0049189,tt0054452"

输出JSON行样例

{"nconst":"nm0000001","primaryName":"Fred Astaire","birthYear":1899,"deathYear":1987,"primaryProfession":"soundtrack,actor,miscellaneous","knownForTitles":"tt0072308,tt0043044,tt0050419,tt0053137"}
{"nconst":"nm0000002","primaryName":"Lauren Bacall","birthYear":1924,"deathYear":2014,"primaryProfession":"actress,soundtrack","knownForTitles":"tt0071877,tt0038355,tt0117057,tt0037382"}
{"nconst":"nm0000003","primaryName":"Brigitte Bardot","birthYear":1934,"deathYear":null,"primaryProfession":"actress,soundtrack,producer","knownForTitles":"tt0057345,tt0059956,tt0049189,tt0054452"}

原有代码错误点

  • 同一文件句柄同时初始化csv.DictReader和csv.reader,两个读取器共用文件指针,会导致读取内容错位、表头获取异常
  • 循环语法不符合Python规则:for i,Value in range(len(keys)), csv_data无法正确并行遍历键列表和行数据,属于逻辑语法错误
  • 缺失要求的类型转换逻辑:未处理\N空值转null的规则,也未对年份字段做数字类型转换
  • 重复实现表头映射逻辑:csv.DictReader本身内置了表头到行字段的映射能力,不需要手动读取表头再循环赋值

完整实现代码

import csv
import json
import os

def myfunccsvtojson(csv_path):
    # 自动生成输出JSON路径,无硬编码,适配任意输入文件名
    file_root, _ = os.path.splitext(csv_path)
    json_output_path = f"{file_root}.json"

    # 流式读写,逐行处理,不加载全量数据到内存
    with open(csv_path, "r", encoding="utf-8") as csv_f, \
         open(json_output_path, "w", encoding="utf-8") as json_f:
        # DictReader自动将第一行表头作为每行字典的键,无需手动映射
        reader = csv.DictReader(csv_f)
        for row in reader:
            processed_row = {}
            for key, value in row.items():
                # \N空值转None,JSON序列化后自动为null
                if value == r"\N":
                    processed_row[key] = None
                # 字段名含Year的年份字段转整数类型
                elif "Year" in key and value.isdigit():
                    processed_row[key] = int(value)
                # 其余字段保留原字符串值,带引号的逗号内容会被csv模块自动正确解析
                else:
                    processed_row[key] = value
            # 逐行写入JSON,符合行分隔格式要求
            json_f.write(json.dumps(processed_row, ensure_ascii=False) + "\n")

逻辑说明

  • 文件名生成通过os.path.splitext拆分输入路径的文件名和后缀,直接替换后缀为.json,支持任意合法文件路径输入
  • 全程采用逐行读取、逐行转换、逐行写入的流式处理逻辑,内存占用恒定,不受CSV文件大小影响,输出的行分隔JSON文件天然支持流式读取
  • DictReader自动跳过表头行,仅遍历处理数据行,10000行输入(含表头)会恰好输出9999行JSON数据
  • 值转换逻辑完全匹配规则:带双引号的逗号分隔字段(如职业列表、作品ID列表)会被csv模块自动识别为单个字段,不会错误拆分

内容的提问来源于stack exchange,提问作者sumanth shetty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 13:57:17