You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多进程优化咨询:批量大体积JSON与CSV互转

Optimizing JSON-to-CSV Conversion with Multi-Processing for Large Files

Got it, let's get your multi-processing setup sorted out—with 16 cores and thousands of 300MB JSON files, we can absolutely cut down your processing time by leveraging parallelism. First, let's break down what we need to adjust in your original code.

Key Observations on Your Original Code

  • Your current script processes files one at a time, which leaves most of your 16 cores sitting idle.
  • You're already handling each JSON file line-by-line (smart move for large files, keeps memory usage low), so we can keep that efficient pattern while parallelizing across files.
  • The broad except clause makes it hard to debug specific errors—we'll tighten that up to catch only JSON parsing issues, so you know exactly what's going wrong when a line fails.

Multi-Processing Solution Using multiprocessing.Pool

We'll wrap your file conversion logic into a reusable function, then use a process pool to distribute the work across your CPU cores. Here's the optimized code:

import json
import csv
import os
from multiprocessing import Pool, cpu_count

def convert_json_to_csv(json_file):
    """Convert a single line-delimited JSON file to CSV"""
    try:
        # Use absolute paths to avoid working directory conflicts in multi-process mode
        json_path = os.path.abspath(json_file)
        partial_name = os.path.splitext(json_file)[0]
        csv_path = f"{partial_name}.csv"

        with open(csv_path, 'w', newline='', encoding='utf-8') as output_file:
            writer = None
            line_num = 1
            with open(json_path, 'r', encoding='utf-8') as file_handle:
                for line in file_handle:
                    try:
                        # Keep floats as strings like your original code
                        data = json.loads(line, parse_float=str)
                    except json.JSONDecodeError as e:
                        print(f"Failed to parse line {line_num} in {json_file}: {str(e)}")
                        line_num += 1
                        continue

                    # Write header on first valid JSON line
                    if writer is None:
                        header = data.keys()
                        writer = csv.writer(output_file)
                        writer.writerow(header)

                    writer.writerow(data.values())
                    line_num += 1
        print(f"Successfully converted {json_file} to {csv_path}")
    except Exception as e:
        print(f"Unexpected error processing {json_file}: {str(e)}")

if __name__ == "__main__":
    # Set your target directory
    target_dir = '/stagingData/Scripts/test'
    os.chdir(target_dir)

    # Filter only JSON files (case-insensitive)
    json_files = [f for f in os.listdir(target_dir) if f.lower().endswith('.json')]

    # Use all available CPU cores (or set to 16 explicitly if needed)
    num_processes = cpu_count()  # This will return 16 on your server
    print(f"Starting conversion with {num_processes} worker processes...")

    # Create a process pool and distribute files to workers
    with Pool(num_processes) as pool:
        pool.map(convert_json_to_csv, json_files)

    print("All conversions completed!")

What We Changed & Why

  • Wrapped conversion logic into a function: Each process handles one file independently, eliminating shared state issues that can cause bugs in multi-process setups.
  • Absolute file paths: Multi-process environments can have inconsistent working directories, so absolute paths prevent "file not found" errors.
  • Targeted exception handling: Now we catch json.JSONDecodeError specifically for parsing issues, plus a top-level exception to handle file IO or other unexpected problems gracefully.
  • Explicit encoding: Adding encoding='utf-8' avoids platform-specific encoding glitches that can corrupt your CSV output.
  • Process pool setup: Pool(num_processes) creates a pool of worker processes, and pool.map distributes your list of JSON files across them. This ensures all 16 cores are put to work simultaneously.

Extra Tips for Large-Scale Processing

  • If disk IO becomes a bottleneck (you notice processes waiting on read/write), consider splitting your files across multiple physical disks to spread the load.
  • For more flexibility, you could use concurrent.futures.ProcessPoolExecutor instead of multiprocessing.Pool—it has a similar API but offers better support for async task tracking.
  • Keep monitoring system resources while running—since you're already using line-by-line parsing, memory usage should stay low, but if you hit issues, you can adjust the number of processes (e.g., drop to 12) to free up resources.

内容的提问来源于stack exchange,提问作者santosh kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:25:38