如何在Python中将含嵌套元素的NDJSON转CSV并拆分大文件
解决方案
合并拆分与CSV转换功能,同时处理嵌套字段扁平化。核心步骤:
- 逐行读取NDJSON文件,避免内存溢出(处理大文件)
- 编写扁平化函数,将嵌套字典转为扁平键值对(如
theme_topic.w1_balanced.label) - 拆分时直接生成CSV文件,无需先保存NDJSON拆分文件
import json import csv from typing import Dict, Any def flatten_dict(nested_dict: Dict[str, Any], parent_key: str = '', sep: str = '.') -> Dict[str, Any]: """扁平化嵌套字典,键用sep连接""" items = [] for k, v in nested_dict.items(): new_key = f"{parent_key}{sep}{k}" if parent_key else k if isinstance(v, dict) and v: items.extend(flatten_dict(v, new_key, sep=sep).items()) else: # 处理空字典或非字典类型,空值设为None items.append((new_key, v if v is not None else None)) return dict(items) def split_and_convert_to_csv(input_file: str, lines_per_file: int): file_count = 0 line_count = 0 rows = [] headers = set() with open(input_file, 'r', encoding="utf8") as infile: for line in infile: line = line.strip() if not line: continue # 解析单条NDJSON行 try: tweet = json.loads(line) # 处理每行是数组包裹单个对象的情况 if isinstance(tweet, list) and len(tweet) > 0: tweet = tweet[0] flat_tweet = flatten_dict(tweet) rows.append(flat_tweet) # 更新表头集合 headers.update(flat_tweet.keys()) line_count += 1 if line_count == lines_per_file: # 生成CSV文件 csv_filename = f'1mio_split_{file_count}.csv' with open(csv_filename, 'w', newline='', encoding="utf8") as csvfile: writer = csv.DictWriter(csvfile, fieldnames=sorted(headers)) writer.writeheader() writer.writerows(rows) print(f"已生成文件: {csv_filename}") # 重置计数器和缓存 file_count += 1 line_count = 0 rows = [] headers = set() except json.JSONDecodeError as e: print(f"解析行失败: {e}, 行内容: {line[:100]}...") continue # 处理剩余行 if rows: csv_filename = f'1mio_split_{file_count}.csv' with open(csv_filename, 'w', newline='', encoding="utf8") as csvfile: writer = csv.DictWriter(csvfile, fieldnames=sorted(headers)) writer.writeheader() writer.writerows(rows) print(f"已生成文件: {csv_filename}") # 用户输入 input_file = input("请输入大文件路径: ") lines_per_file = int(input("每个文件包含多少行推文?: ")) # 执行拆分转换 split_and_convert_to_csv(input_file, lines_per_file) print("拆分与转换完成!")
代码说明
- 扁平化函数:
flatten_dict递归处理嵌套字典,将嵌套键转为父键.子键的形式,比如theme_topic.w1_balanced.confidence - NDJSON解析:逐行读取并解析,处理每行是数组包裹单个对象的情况
- 内存优化:每攒够指定行数就写入CSV并清空缓存,避免加载大文件到内存
- 错误处理:跳过解析失败的行,打印错误信息不中断流程
内容的提问来源于stack exchange,提问作者scraper34863
相关产品推荐
相关产品推荐

