You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从大文本文件提取JSON字符串:OpenLibrary Editions Dump处理需求

Got it, let's tackle extracting the JSON strings from that 5GB OpenLibrary Editions Dump—super common task when working with these dumps, so I've got a couple of solid methods for you depending on what you need.

Methods to Extract JSON from OpenLibrary Editions Dump

First, let's recap the line structure of the dump: each row starts with 4 non-JSON fields (/type/edition, book key, version, timestamp), followed by the JSON object we care about. Our goal is to strip off those first 4 fields and keep only the JSON.

1. Command-Line Tools (Fast for Large 5GB Files)

If you just need a quick extraction without extra processing, command-line tools are ideal—they're memory-efficient and handle big files way faster than loading everything into a script.

Option 1: Awk

This one-liner replaces everything from the start of the line up to the first { with just {, leaving you with the full JSON string:

awk '{sub(/^.*{/,"{"); print}' your_dump_file > extracted_editions.json

Option 2: Grep (With PCRE Support)

Using grep's Perl-compatible regex, we can skip the first 4 whitespace-separated fields and capture everything after:

grep -oP '^\S+\s+\S+\s+\S+\s+\S+\s+\K.*' your_dump_file > extracted_editions.json

The \K tells grep to discard everything matched before it, so we only get the JSON portion.

2. Python Script (For Validation & Post-Processing)

If you need to validate the JSON (to skip bad lines) or plan to process the data further, a Python script is the way to go. It reads the file line-by-line (so it won't choke on the 5GB size) and can handle validation:

import json

# Update these paths to match your files
INPUT_DUMP = "path/to/your/openlibrary_dump.txt"
OUTPUT_JSON = "extracted_valid_editions.json"

with open(INPUT_DUMP, 'r', encoding='utf-8') as infile, \
     open(OUTPUT_JSON, 'w', encoding='utf-8') as outfile:

    for line_num, line in enumerate(infile, 1):
        line = line.strip()
        if not line:
            continue
        
        # Split into first 4 fields + JSON string
        parts = line.split(maxsplit=4)
        if len(parts) < 5:
            print(f"Skipping malformed line {line_num}: missing JSON portion")
            continue
        
        json_str = parts[4]
        # Optional: Validate JSON to ensure it's valid
        try:
            json_data = json.loads(json_str)
            # Write as a JSON Lines object (one per line)
            json.dump(json_data, outfile)
            outfile.write('\n')
        except json.JSONDecodeError as e:
            print(f"Invalid JSON in line {line_num}: {str(e)}")

This script will:

  • Skip empty lines
  • Flag lines that don't have a JSON portion
  • Validate each JSON string and skip invalid ones
  • Output valid JSON objects in JSON Lines format (easy to parse later)

Bonus: JSON Array Output

If you need all JSON objects in a single array (instead of one per line), modify the script slightly:

import json

INPUT_DUMP = "path/to/your/openlibrary_dump.txt"
OUTPUT_JSON = "extracted_editions_array.json"

json_objects = []

with open(INPUT_DUMP, 'r', encoding='utf-8') as infile:
    for line_num, line in enumerate(infile, 1):
        line = line.strip()
        if not line:
            continue
        
        parts = line.split(maxsplit=4)
        if len(parts) < 5:
            continue
        
        json_str = parts[4]
        try:
            json_objects.append(json.loads(json_str))
        except json.JSONDecodeError:
            continue

with open(OUTPUT_JSON, 'w', encoding='utf-8') as outfile:
    json.dump(json_objects, outfile, indent=2)

Note: This will load all valid JSON into memory—for 5GB of data, this might require a lot of RAM, so only use this if you have enough memory available.

内容的提问来源于stack exchange,提问作者Tom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 11:11:36