Python多进程优化咨询:批量大体积JSON与CSV互转
Optimizing JSON-to-CSV Conversion with Multi-Processing for Large Files
Got it, let's get your multi-processing setup sorted out—with 16 cores and thousands of 300MB JSON files, we can absolutely cut down your processing time by leveraging parallelism. First, let's break down what we need to adjust in your original code.
Key Observations on Your Original Code
- Your current script processes files one at a time, which leaves most of your 16 cores sitting idle.
- You're already handling each JSON file line-by-line (smart move for large files, keeps memory usage low), so we can keep that efficient pattern while parallelizing across files.
- The broad
exceptclause makes it hard to debug specific errors—we'll tighten that up to catch only JSON parsing issues, so you know exactly what's going wrong when a line fails.
Multi-Processing Solution Using multiprocessing.Pool
We'll wrap your file conversion logic into a reusable function, then use a process pool to distribute the work across your CPU cores. Here's the optimized code:
import json import csv import os from multiprocessing import Pool, cpu_count def convert_json_to_csv(json_file): """Convert a single line-delimited JSON file to CSV""" try: # Use absolute paths to avoid working directory conflicts in multi-process mode json_path = os.path.abspath(json_file) partial_name = os.path.splitext(json_file)[0] csv_path = f"{partial_name}.csv" with open(csv_path, 'w', newline='', encoding='utf-8') as output_file: writer = None line_num = 1 with open(json_path, 'r', encoding='utf-8') as file_handle: for line in file_handle: try: # Keep floats as strings like your original code data = json.loads(line, parse_float=str) except json.JSONDecodeError as e: print(f"Failed to parse line {line_num} in {json_file}: {str(e)}") line_num += 1 continue # Write header on first valid JSON line if writer is None: header = data.keys() writer = csv.writer(output_file) writer.writerow(header) writer.writerow(data.values()) line_num += 1 print(f"Successfully converted {json_file} to {csv_path}") except Exception as e: print(f"Unexpected error processing {json_file}: {str(e)}") if __name__ == "__main__": # Set your target directory target_dir = '/stagingData/Scripts/test' os.chdir(target_dir) # Filter only JSON files (case-insensitive) json_files = [f for f in os.listdir(target_dir) if f.lower().endswith('.json')] # Use all available CPU cores (or set to 16 explicitly if needed) num_processes = cpu_count() # This will return 16 on your server print(f"Starting conversion with {num_processes} worker processes...") # Create a process pool and distribute files to workers with Pool(num_processes) as pool: pool.map(convert_json_to_csv, json_files) print("All conversions completed!")
What We Changed & Why
- Wrapped conversion logic into a function: Each process handles one file independently, eliminating shared state issues that can cause bugs in multi-process setups.
- Absolute file paths: Multi-process environments can have inconsistent working directories, so absolute paths prevent "file not found" errors.
- Targeted exception handling: Now we catch
json.JSONDecodeErrorspecifically for parsing issues, plus a top-level exception to handle file IO or other unexpected problems gracefully. - Explicit encoding: Adding
encoding='utf-8'avoids platform-specific encoding glitches that can corrupt your CSV output. - Process pool setup:
Pool(num_processes)creates a pool of worker processes, andpool.mapdistributes your list of JSON files across them. This ensures all 16 cores are put to work simultaneously.
Extra Tips for Large-Scale Processing
- If disk IO becomes a bottleneck (you notice processes waiting on read/write), consider splitting your files across multiple physical disks to spread the load.
- For more flexibility, you could use
concurrent.futures.ProcessPoolExecutorinstead ofmultiprocessing.Pool—it has a similar API but offers better support for async task tracking. - Keep monitoring system resources while running—since you're already using line-by-line parsing, memory usage should stay low, but if you hit issues, you can adjust the number of processes (e.g., drop to 12) to free up resources.
内容的提问来源于stack exchange,提问作者santosh kumar
相关产品推荐
相关产品推荐

