You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

非标准JSON转CSV:700MB无分隔符JSON数据转换方法问询

Convert JSON Lines (Non-Standard JSON) to CSV Efficiently for Large Files

Got it, let's break this down—your data is actually in JSON Lines (JSONL) format, where each line is a standalone valid JSON object (no commas between them). The issue with "regular" methods is probably because tools expecting a single valid JSON array choke on this, and loading a 700MB file fully into memory causes performance hits or crashes. Here are two reliable, efficient solutions:

Python Approach (Memory-Friendly)

This reads the file line by line, so it won't hog memory even with large datasets. We'll use Python's built-in csv and json modules—no extra dependencies needed:

import json
import csv

# Define your desired CSV headers
csv_headers = ["hash", "block_timestamp", "addresses"]

with open("input.json", "r") as infile, open("output.csv", "w", newline="") as outfile:
    writer = csv.DictWriter(outfile, fieldnames=csv_headers)
    writer.writeheader()
    
    for line in infile:
        # Skip any empty lines that might be in the file
        cleaned_line = line.strip()
        if not cleaned_line:
            continue
        # Parse each individual JSON object
        entry = json.loads(cleaned_line)
        # Convert the addresses array to a comma-separated string (adjust if you want to keep the array format)
        entry["addresses"] = ", ".join(entry["addresses"])
        writer.writerow(entry)

Quick Notes:

  • If you want to keep the addresses field in its original JSON array format (like ["3E17PiWGJqP8945KRZHuPdsFSU59othGEQ"]), just remove the entry["addresses"] = ", ".join(...) line.
  • The newline="" parameter fixes extra blank lines in CSV outputs on Windows.

BASH + jq Approach (Fast, No Python Required)

If you prefer command-line tools, jq is made for handling JSONL efficiently. First, make sure jq is installed (most Linux distros include it; macOS users can get it via Homebrew, Windows via Chocolatey).

Run this command:

jq -r '["hash", "block_timestamp", "addresses"] as $headers | $headers, (.addresses | join(", ")) as $addr | [.hash, .block_timestamp, $addr] | @csv' input.json > output.csv

What This Does:

  • -r outputs raw strings instead of JSON-encoded values.
  • We first define the CSV headers, then process each line:
    • Convert the addresses array to a comma-separated string with join(", ").
    • Assemble the fields into an array, then convert to proper CSV format with @csv.
  • To keep the addresses array as JSON, replace (.addresses | join(", ")) as $addr | [.hash, .block_timestamp, $addr] with [.hash, .block_timestamp, .addresses].

Why This Works for 700MB Files:

Both methods process the file line by line, so they never load the entire dataset into memory. This avoids the slowdowns or crashes you might get with tools that try to parse the whole file at once.

内容的提问来源于stack exchange,提问作者Warrior..

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 15:03:12