使用Pandas转JSON为CSV时触发JSONDecodeError,请求排查解决
First off, let's break down why your first file worked but the rest are throwing JSONDecodeError: Expecting value—even if the files look identical, there are subtle, easy-to-miss issues that can trip up strict JSON parsers like the one Pandas uses under the hood. Here's how to diagnose and fix this:
Step 1: Validate the Problematic JSON Files
First, confirm if the files are actually valid JSON. The "Expecting value" error usually means the parser hit something it didn't expect (like a missing comma, extra character, or malformed structure).
Use this quick Python script to check each file and pinpoint the error:
import json def check_json_validity(file_path): try: with open(file_path, 'r', encoding='utf-8') as f: json.load(f) print(f"✅ {file_path} is valid JSON") except json.JSONDecodeError as e: print(f"❌ {file_path} has an error: {e}") # Print the line where the error occurred to debug with open(file_path, 'r', encoding='utf-8') as f: lines = f.readlines() if e.lineno <= len(lines): print(f"Error near line {e.lineno}: {lines[e.lineno-1].strip()}") # Replace with your problematic file path check_json_validity(r'path\to\your\problem_file.json')
Common Issues & Fixes
1. Hidden Trailing Commas
JSON doesn't allow trailing commas after the last element in an array or object (e.g., "ted" : "yes" }, ] }). Some tools or generators accidentally add these, and while some parsers ignore them, Pandas' strict parser won't.
Fix: Use a more lenient JSON parser like json5 (install with pip install json5) to load the file, then convert to a DataFrame:
import json5 import pandas as pd # Load the leniently parsed JSON with open(r'path\to\problem_file.json', 'r', encoding='utf-8') as f: data = json5.load(f) # Convert the 'data' array to a DataFrame df = pd.DataFrame(data['data']) # Optional: Flatten nested fields (like 'names' and 'active') for CSV df['names'] = df['names'].apply(lambda x: ','.join(x)) # Turn list into comma-separated string df = pd.concat([df.drop('active', axis=1), df['active'].apply(pd.Series)], axis=1) # Expand 'active' into columns # Save to CSV df.to_csv(r'path\to\output.csv', index=False)
2. Encoding Mismatches
Your first file might be UTF-8, but others could be encoded with BOM (UTF-8-SIG) or another format like GBK. The parser can't interpret the hidden encoding markers, leading to errors.
Fix: Specify the correct encoding when reading the file:
import pandas as pd # Try UTF-8 with BOM first df = pd.read_json(r'path\to\problem_file.json', encoding='utf-8-sig') # If that fails, try other encodings like 'gbk' for Chinese text # df = pd.read_json(r'path\to\problem_file.json', encoding='gbk')
3. Incorrect orient Parameter
Your original code uses orient="split", which is designed for JSON structured like {"columns": [...], "index": [...], "data": [...]}. Your sample JSON has a data array of objects, so this parameter might have worked by accident for the first file but fails for others.
Fix: Load the JSON manually first, then convert the data array to a DataFrame (this is more reliable):
import json import pandas as pd with open(r'path\to\json.json', 'r', encoding='utf-8') as f: json_data = json.load(f) # Convert the 'data' array to a DataFrame df = pd.DataFrame(json_data['data']) df.to_csv(r'path\to\output.csv', index=False)
4. Multiple JSON Objects in One File
If your files have one JSON object per line (instead of a single nested structure), orient="split" won't work.
Fix: Use lines=True to read line-delimited JSON:
df = pd.read_json(r'path\to\problem_file.json', lines=True)
Final Tips
- Always validate JSON files before processing to catch issues early.
- Avoid relying on
orientparameters unless you're sure the JSON matches that structure—loading the JSON manually gives you more control.
内容的提问来源于stack exchange,提问作者nos codemos

