如何在Python中读取单文件内的多个独立JSON数据?
Got it, let's tackle this problem step by step. You're dealing with that annoying scenario where each individual JSON object is valid, but the whole file isn't a proper JSON array—super common when working with log files or bulk-exported data! Here are two reliable approaches to get the exact result you want:
Method 1: Base Python + Pandas (Handles Nested JSON & Unformatted Files)
This method is the most robust, especially if your JSON objects are squished together without line breaks or have nested JSON structures. We'll use Python's built-in json module to safely parse each complete JSON object one by one:
import json import pandas as pd # Replace with your file path file_path = '~/Desktop/data.json' # Read the entire file content with open(file_path, 'r') as f: content = f.read() # Use JSONDecoder to parse each valid JSON object sequentially decoder = json.JSONDecoder() current_pos = 0 json_objects = [] while current_pos < len(content): try: # Parse the next valid JSON object and get the end position obj, current_pos = decoder.raw_decode(content, current_pos) json_objects.append(obj) # Skip any whitespace between objects (spaces, newlines, tabs) while current_pos < len(content) and content[current_pos].isspace(): current_pos += 1 except json.JSONDecodeError: # If we hit invalid content, skip one character and keep going (adjust this if needed) current_pos += 1 # Convert the list of JSON objects (Python dicts) to a DataFrame df = pd.DataFrame({"jsons": json_objects})
Notes:
- This correctly handles nested JSON because
raw_decodeidentifies the complete structure of each JSON object, so it won't get tripped up by nested{}pairs. - If you want the cells to hold JSON strings instead of Python dictionaries, replace
json_objects.append(obj)withjson_objects.append(json.dumps(obj)).
Method 2: Quick Pandas Adjustment (If lines=True Already Works)
If your file does have one JSON object per line (even if they're not comma-separated), you can use your existing pd.read_json(..., lines=True) approach and then merge the split columns back into a single JSON object per row:
import pandas as pd # First read the file with lines=True (splits into columns) df_split = pd.read_json('~/Desktop/data.json', lines=True) # Convert each row back to a full JSON object/dictionary df = pd.DataFrame({"jsons": df_split.apply(lambda row: row.to_dict(), axis=1)})
Notes:
- If you want JSON strings instead of dictionaries, use
row.to_json()instead ofrow.to_dict(). - This is faster and simpler, but only works if each JSON object is on its own line (which
lines=Truerelies on).
Either method will give you a DataFrame where each row's jsons column holds the full, original JSON object—no more split columns messing up your structure!
内容的提问来源于stack exchange,提问作者Outcast

