非标准JSON转CSV:700MB无分隔符JSON数据转换方法问询
Got it, let's break this down—your data is actually in JSON Lines (JSONL) format, where each line is a standalone valid JSON object (no commas between them). The issue with "regular" methods is probably because tools expecting a single valid JSON array choke on this, and loading a 700MB file fully into memory causes performance hits or crashes. Here are two reliable, efficient solutions:
Python Approach (Memory-Friendly)
This reads the file line by line, so it won't hog memory even with large datasets. We'll use Python's built-in csv and json modules—no extra dependencies needed:
import json import csv # Define your desired CSV headers csv_headers = ["hash", "block_timestamp", "addresses"] with open("input.json", "r") as infile, open("output.csv", "w", newline="") as outfile: writer = csv.DictWriter(outfile, fieldnames=csv_headers) writer.writeheader() for line in infile: # Skip any empty lines that might be in the file cleaned_line = line.strip() if not cleaned_line: continue # Parse each individual JSON object entry = json.loads(cleaned_line) # Convert the addresses array to a comma-separated string (adjust if you want to keep the array format) entry["addresses"] = ", ".join(entry["addresses"]) writer.writerow(entry)
Quick Notes:
- If you want to keep the
addressesfield in its original JSON array format (like["3E17PiWGJqP8945KRZHuPdsFSU59othGEQ"]), just remove theentry["addresses"] = ", ".join(...)line. - The
newline=""parameter fixes extra blank lines in CSV outputs on Windows.
BASH + jq Approach (Fast, No Python Required)
If you prefer command-line tools, jq is made for handling JSONL efficiently. First, make sure jq is installed (most Linux distros include it; macOS users can get it via Homebrew, Windows via Chocolatey).
Run this command:
jq -r '["hash", "block_timestamp", "addresses"] as $headers | $headers, (.addresses | join(", ")) as $addr | [.hash, .block_timestamp, $addr] | @csv' input.json > output.csv
What This Does:
-routputs raw strings instead of JSON-encoded values.- We first define the CSV headers, then process each line:
- Convert the
addressesarray to a comma-separated string withjoin(", "). - Assemble the fields into an array, then convert to proper CSV format with
@csv.
- Convert the
- To keep the
addressesarray as JSON, replace(.addresses | join(", ")) as $addr | [.hash, .block_timestamp, $addr]with[.hash, .block_timestamp, .addresses].
Why This Works for 700MB Files:
Both methods process the file line by line, so they never load the entire dataset into memory. This avoids the slowdowns or crashes you might get with tools that try to parse the whole file at once.
内容的提问来源于stack exchange,提问作者Warrior..

