You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中将含嵌套元素的NDJSON转CSV并拆分大文件

解决方案

合并拆分与CSV转换功能,同时处理嵌套字段扁平化。核心步骤:

  1. 逐行读取NDJSON文件,避免内存溢出(处理大文件)
  2. 编写扁平化函数,将嵌套字典转为扁平键值对(如theme_topic.w1_balanced.label)
  3. 拆分时直接生成CSV文件,无需先保存NDJSON拆分文件
import json
import csv
from typing import Dict, Any

def flatten_dict(nested_dict: Dict[str, Any], parent_key: str = '', sep: str = '.') -> Dict[str, Any]:
    """扁平化嵌套字典,键用sep连接"""
    items = []
    for k, v in nested_dict.items():
        new_key = f"{parent_key}{sep}{k}" if parent_key else k
        if isinstance(v, dict) and v:
            items.extend(flatten_dict(v, new_key, sep=sep).items())
        else:
            # 处理空字典或非字典类型,空值设为None
            items.append((new_key, v if v is not None else None))
    return dict(items)

def split_and_convert_to_csv(input_file: str, lines_per_file: int):
    file_count = 0
    line_count = 0
    rows = []
    headers = set()

    with open(input_file, 'r', encoding="utf8") as infile:
        for line in infile:
            line = line.strip()
            if not line:
                continue
            # 解析单条NDJSON行
            try:
                tweet = json.loads(line)
                # 处理每行是数组包裹单个对象的情况
                if isinstance(tweet, list) and len(tweet) > 0:
                    tweet = tweet[0]
                flat_tweet = flatten_dict(tweet)
                rows.append(flat_tweet)
                # 更新表头集合
                headers.update(flat_tweet.keys())
                line_count += 1

                if line_count == lines_per_file:
                    # 生成CSV文件
                    csv_filename = f'1mio_split_{file_count}.csv'
                    with open(csv_filename, 'w', newline='', encoding="utf8") as csvfile:
                        writer = csv.DictWriter(csvfile, fieldnames=sorted(headers))
                        writer.writeheader()
                        writer.writerows(rows)
                    print(f"已生成文件: {csv_filename}")
                    # 重置计数器和缓存
                    file_count += 1
                    line_count = 0
                    rows = []
                    headers = set()
            except json.JSONDecodeError as e:
                print(f"解析行失败: {e}, 行内容: {line[:100]}...")
                continue
        # 处理剩余行
        if rows:
            csv_filename = f'1mio_split_{file_count}.csv'
            with open(csv_filename, 'w', newline='', encoding="utf8") as csvfile:
                writer = csv.DictWriter(csvfile, fieldnames=sorted(headers))
                writer.writeheader()
                writer.writerows(rows)
            print(f"已生成文件: {csv_filename}")

# 用户输入
input_file = input("请输入大文件路径: ")
lines_per_file = int(input("每个文件包含多少行推文?: "))

# 执行拆分转换
split_and_convert_to_csv(input_file, lines_per_file)
print("拆分与转换完成!")

代码说明

  • 扁平化函数:flatten_dict递归处理嵌套字典,将嵌套键转为父键.子键的形式,比如theme_topic.w1_balanced.confidence
  • NDJSON解析:逐行读取并解析,处理每行是数组包裹单个对象的情况
  • 内存优化:每攒够指定行数就写入CSV并清空缓存,避免加载大文件到内存
  • 错误处理:跳过解析失败的行,打印错误信息不中断流程

内容的提问来源于stack exchange,提问作者scraper34863

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 07:57:03