You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中读取单文件内的多个独立JSON数据?

Solution: Get a DataFrame with Complete JSON Objects per Cell

Got it, let's tackle this problem step by step. You're dealing with that annoying scenario where each individual JSON object is valid, but the whole file isn't a proper JSON array—super common when working with log files or bulk-exported data! Here are two reliable approaches to get the exact result you want:

Method 1: Base Python + Pandas (Handles Nested JSON & Unformatted Files)

This method is the most robust, especially if your JSON objects are squished together without line breaks or have nested JSON structures. We'll use Python's built-in json module to safely parse each complete JSON object one by one:

import json
import pandas as pd

# Replace with your file path
file_path = '~/Desktop/data.json'

# Read the entire file content
with open(file_path, 'r') as f:
    content = f.read()

# Use JSONDecoder to parse each valid JSON object sequentially
decoder = json.JSONDecoder()
current_pos = 0
json_objects = []

while current_pos < len(content):
    try:
        # Parse the next valid JSON object and get the end position
        obj, current_pos = decoder.raw_decode(content, current_pos)
        json_objects.append(obj)
        
        # Skip any whitespace between objects (spaces, newlines, tabs)
        while current_pos < len(content) and content[current_pos].isspace():
            current_pos += 1
    except json.JSONDecodeError:
        # If we hit invalid content, skip one character and keep going (adjust this if needed)
        current_pos += 1

# Convert the list of JSON objects (Python dicts) to a DataFrame
df = pd.DataFrame({"jsons": json_objects})

Notes:

  • This correctly handles nested JSON because raw_decode identifies the complete structure of each JSON object, so it won't get tripped up by nested {} pairs.
  • If you want the cells to hold JSON strings instead of Python dictionaries, replace json_objects.append(obj) with json_objects.append(json.dumps(obj)).

Method 2: Quick Pandas Adjustment (If lines=True Already Works)

If your file does have one JSON object per line (even if they're not comma-separated), you can use your existing pd.read_json(..., lines=True) approach and then merge the split columns back into a single JSON object per row:

import pandas as pd

# First read the file with lines=True (splits into columns)
df_split = pd.read_json('~/Desktop/data.json', lines=True)

# Convert each row back to a full JSON object/dictionary
df = pd.DataFrame({"jsons": df_split.apply(lambda row: row.to_dict(), axis=1)})

Notes:

  • If you want JSON strings instead of dictionaries, use row.to_json() instead of row.to_dict().
  • This is faster and simpler, but only works if each JSON object is on its own line (which lines=True relies on).

Either method will give you a DataFrame where each row's jsons column holds the full, original JSON object—no more split columns messing up your structure!

内容的提问来源于stack exchange,提问作者Outcast

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:05:43