You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python将大体积corpus.json文件转换为TSV格式

Convert Large corpus.json to TSV with Python

Hey there! Converting a massive corpus.json to TSV in Python is totally manageable—we just need to prioritize efficient memory usage since the file is so large. Let’s break this down based on your JSON structure, where top-level numeric keys (like "0") point to objects containing a metadata field with id, title, tags (an array), and description.

Key Priority for Large Files

We won’t load the entire JSON into memory at once—that’s a surefire way to hit memory limits. Instead, we’ll use streaming or line-by-line parsing to process data incrementally.

Method 1: Stream Parsing with ijson (Best for Single Large JSON File)

The ijson library lets us parse JSON piece by piece, which is ideal for files that would crash a standard json.load() call.

Step 1: Install the Library

pip install ijson

Step 2: Python Code

import ijson
import csv

# Define the fields we want to extract from metadata
TARGET_FIELDS = ["id", "title", "tags", "description"]

with open("corpus.json", "rb") as json_in, open("corpus.tsv", "w", newline="", encoding="utf-8") as tsv_out:
    # Set up TSV writer with tab delimiter
    tsv_writer = csv.writer(tsv_out, delimiter="\t")
    # Write header row
    tsv_writer.writerow(TARGET_FIELDS)

    # Stream through every metadata object in the JSON
    for metadata in ijson.items(json_in, "item.metadata"):
        # Convert tags array to a comma-separated string (TSV doesn't handle nested arrays)
        formatted_tags = ", ".join(metadata.get("tags", []))
        
        # Extract fields (use empty string if a field is missing)
        row_data = [
            metadata.get("id", ""),
            metadata.get("title", ""),
            formatted_tags,
            metadata.get("description", "")
        ]

        # Clean tabs/newlines to avoid breaking TSV structure
        cleaned_row = [str(cell).replace("\t", " ").replace("\n", " ") for cell in row_data]
        tsv_writer.writerow(cleaned_row)

How This Works:

  • ijson.items(json_in, "item.metadata") targets every metadata object nested under the top-level numeric keys.
  • We convert the tags array to a string to keep the TSV format clean.
  • Cleaning tabs and newlines prevents misalignment in the output file.

Method 2: Line-by-Line Parsing (For JSON Lines Format)

If your corpus.json uses one JSON object per line (JSON Lines format), you can use Python’s built-in modules without extra installs:

import json
import csv

TARGET_FIELDS = ["id", "title", "tags", "description"]

with open("corpus.json", "r", encoding="utf-8") as json_in, open("corpus.tsv", "w", newline="", encoding="utf-8") as tsv_out:
    tsv_writer = csv.writer(tsv_out, delimiter="\t")
    tsv_writer.writerow(TARGET_FIELDS)

    for line in json_in:
        line = line.strip()
        if not line:
            continue  # Skip empty lines
        
        try:
            # Parse the line as JSON
            line_data = json.loads(line)
            # Iterate through top-level numeric keys
            for item in line_data.values():
                metadata = item.get("metadata", {})
                formatted_tags = ", ".join(metadata.get("tags", []))
                
                row_data = [
                    metadata.get("id", ""),
                    metadata.get("title", ""),
                    formatted_tags,
                    metadata.get("description", "")
                ]
                
                cleaned_row = [str(cell).replace("\t", " ").replace("\n", " ") for cell in row_data]
                tsv_writer.writerow(cleaned_row)
        except json.JSONDecodeError as e:
            print(f"Skipping invalid line: {line[:50]}... Error: {e}")

Important Notes

  • Encoding: Always specify encoding="utf-8" to handle special characters correctly.
  • Missing Fields: Using .get(field, "") prevents KeyError if some metadata objects lack certain fields.
  • Performance: Both methods use minimal memory by processing data incrementally—critical for large files.

内容的提问来源于stack exchange,提问作者Nidhi Agarwal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:52:51