使用Python将大体积corpus.json文件转换为TSV格式
Hey there! Converting a massive corpus.json to TSV in Python is totally manageable—we just need to prioritize efficient memory usage since the file is so large. Let’s break this down based on your JSON structure, where top-level numeric keys (like "0") point to objects containing a metadata field with id, title, tags (an array), and description.
Key Priority for Large Files
We won’t load the entire JSON into memory at once—that’s a surefire way to hit memory limits. Instead, we’ll use streaming or line-by-line parsing to process data incrementally.
Method 1: Stream Parsing with ijson (Best for Single Large JSON File)
The ijson library lets us parse JSON piece by piece, which is ideal for files that would crash a standard json.load() call.
Step 1: Install the Library
pip install ijson
Step 2: Python Code
import ijson import csv # Define the fields we want to extract from metadata TARGET_FIELDS = ["id", "title", "tags", "description"] with open("corpus.json", "rb") as json_in, open("corpus.tsv", "w", newline="", encoding="utf-8") as tsv_out: # Set up TSV writer with tab delimiter tsv_writer = csv.writer(tsv_out, delimiter="\t") # Write header row tsv_writer.writerow(TARGET_FIELDS) # Stream through every metadata object in the JSON for metadata in ijson.items(json_in, "item.metadata"): # Convert tags array to a comma-separated string (TSV doesn't handle nested arrays) formatted_tags = ", ".join(metadata.get("tags", [])) # Extract fields (use empty string if a field is missing) row_data = [ metadata.get("id", ""), metadata.get("title", ""), formatted_tags, metadata.get("description", "") ] # Clean tabs/newlines to avoid breaking TSV structure cleaned_row = [str(cell).replace("\t", " ").replace("\n", " ") for cell in row_data] tsv_writer.writerow(cleaned_row)
How This Works:
ijson.items(json_in, "item.metadata")targets everymetadataobject nested under the top-level numeric keys.- We convert the
tagsarray to a string to keep the TSV format clean. - Cleaning tabs and newlines prevents misalignment in the output file.
Method 2: Line-by-Line Parsing (For JSON Lines Format)
If your corpus.json uses one JSON object per line (JSON Lines format), you can use Python’s built-in modules without extra installs:
import json import csv TARGET_FIELDS = ["id", "title", "tags", "description"] with open("corpus.json", "r", encoding="utf-8") as json_in, open("corpus.tsv", "w", newline="", encoding="utf-8") as tsv_out: tsv_writer = csv.writer(tsv_out, delimiter="\t") tsv_writer.writerow(TARGET_FIELDS) for line in json_in: line = line.strip() if not line: continue # Skip empty lines try: # Parse the line as JSON line_data = json.loads(line) # Iterate through top-level numeric keys for item in line_data.values(): metadata = item.get("metadata", {}) formatted_tags = ", ".join(metadata.get("tags", [])) row_data = [ metadata.get("id", ""), metadata.get("title", ""), formatted_tags, metadata.get("description", "") ] cleaned_row = [str(cell).replace("\t", " ").replace("\n", " ") for cell in row_data] tsv_writer.writerow(cleaned_row) except json.JSONDecodeError as e: print(f"Skipping invalid line: {line[:50]}... Error: {e}")
Important Notes
- Encoding: Always specify
encoding="utf-8"to handle special characters correctly. - Missing Fields: Using
.get(field, "")preventsKeyErrorif somemetadataobjects lack certain fields. - Performance: Both methods use minimal memory by processing data incrementally—critical for large files.
内容的提问来源于stack exchange,提问作者Nidhi Agarwal

