从大文本文件提取JSON字符串:OpenLibrary Editions Dump处理需求
Got it, let's tackle extracting the JSON strings from that 5GB OpenLibrary Editions Dump—super common task when working with these dumps, so I've got a couple of solid methods for you depending on what you need.
First, let's recap the line structure of the dump: each row starts with 4 non-JSON fields (/type/edition, book key, version, timestamp), followed by the JSON object we care about. Our goal is to strip off those first 4 fields and keep only the JSON.
1. Command-Line Tools (Fast for Large 5GB Files)
If you just need a quick extraction without extra processing, command-line tools are ideal—they're memory-efficient and handle big files way faster than loading everything into a script.
Option 1: Awk
This one-liner replaces everything from the start of the line up to the first { with just {, leaving you with the full JSON string:
awk '{sub(/^.*{/,"{"); print}' your_dump_file > extracted_editions.json
Option 2: Grep (With PCRE Support)
Using grep's Perl-compatible regex, we can skip the first 4 whitespace-separated fields and capture everything after:
grep -oP '^\S+\s+\S+\s+\S+\s+\S+\s+\K.*' your_dump_file > extracted_editions.json
The \K tells grep to discard everything matched before it, so we only get the JSON portion.
2. Python Script (For Validation & Post-Processing)
If you need to validate the JSON (to skip bad lines) or plan to process the data further, a Python script is the way to go. It reads the file line-by-line (so it won't choke on the 5GB size) and can handle validation:
import json # Update these paths to match your files INPUT_DUMP = "path/to/your/openlibrary_dump.txt" OUTPUT_JSON = "extracted_valid_editions.json" with open(INPUT_DUMP, 'r', encoding='utf-8') as infile, \ open(OUTPUT_JSON, 'w', encoding='utf-8') as outfile: for line_num, line in enumerate(infile, 1): line = line.strip() if not line: continue # Split into first 4 fields + JSON string parts = line.split(maxsplit=4) if len(parts) < 5: print(f"Skipping malformed line {line_num}: missing JSON portion") continue json_str = parts[4] # Optional: Validate JSON to ensure it's valid try: json_data = json.loads(json_str) # Write as a JSON Lines object (one per line) json.dump(json_data, outfile) outfile.write('\n') except json.JSONDecodeError as e: print(f"Invalid JSON in line {line_num}: {str(e)}")
This script will:
- Skip empty lines
- Flag lines that don't have a JSON portion
- Validate each JSON string and skip invalid ones
- Output valid JSON objects in JSON Lines format (easy to parse later)
Bonus: JSON Array Output
If you need all JSON objects in a single array (instead of one per line), modify the script slightly:
import json INPUT_DUMP = "path/to/your/openlibrary_dump.txt" OUTPUT_JSON = "extracted_editions_array.json" json_objects = [] with open(INPUT_DUMP, 'r', encoding='utf-8') as infile: for line_num, line in enumerate(infile, 1): line = line.strip() if not line: continue parts = line.split(maxsplit=4) if len(parts) < 5: continue json_str = parts[4] try: json_objects.append(json.loads(json_str)) except json.JSONDecodeError: continue with open(OUTPUT_JSON, 'w', encoding='utf-8') as outfile: json.dump(json_objects, outfile, indent=2)
Note: This will load all valid JSON into memory—for 5GB of data, this might require a lot of RAM, so only use this if you have enough memory available.
内容的提问来源于stack exchange,提问作者Tom

