You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:从含lista及数据样例图的嵌套JSON生成按Post ID拆分的Pandas DataFrame

Solution: Extract Text from Nested JSON and Create Per-Post DataFrames

Hey there! Let’s break down how to solve this problem step by step. Since your JSON is highly nested, we’ll need to traverse its structure to pull out text content and map it to each Post ID, then build separate Pandas DataFrames for each post.

Step 1: Import Required Libraries

First, make sure you have the necessary tools ready (json is built-in, but you’ll need pandas installed if you don’t have it already):

import json
import pandas as pd

Step 2: Load Your JSON Data

Start by reading the JSON file into a Python dictionary:

# Replace 'your_nested_data.json' with your actual file path
with open('your_nested_data.json', 'r', encoding='utf-8') as f:
    json_data = json.load(f)

Step 3: Extract Text and Build DataFrames

Case 1: Known Nested Structure

If you know exactly where the Post ID and text are located in your JSON (e.g., under lista -> post -> id and lista -> post -> content -> sections -> text), you can directly traverse those paths. Here’s an example using a structure similar to what you described:

# Initialize a dictionary to store DataFrames (key: Post ID, value: DataFrame)
post_dfs = {}

# Iterate over each entry in the 'lista' array
for item in json_data['lista']:
    # Extract Post ID (adjust this path to match your JSON's actual structure)
    post_id = item['post']['id']
    
    # Extract all text blocks (tweak this path based on where your text lives)
    text_blocks = [section['text'] for section in item['post']['content']['sections']]
    
    # Create a DataFrame for this post with a 'Text' column
    post_df = pd.DataFrame({'Text': text_blocks})
    
    # Add the DataFrame to our dictionary using the Post ID as the key
    post_dfs[post_id] = post_df

Case 2: Deeply Unpredictable Nesting

If text could be nested in varying levels (and you want to capture every piece of text associated with a post), use a recursive function to hunt down all instances of text:

def extract_all_text(nested_element):
    """Recursively pull out all 'text' values from a nested JSON structure."""
    text_list = []
    if isinstance(nested_element, dict):
        for key, value in nested_element.items():
            if key == 'text':
                text_list.append(value)
            else:
                text_list.extend(extract_all_text(value))
    elif isinstance(nested_element, list):
        for item in nested_element:
            text_list.extend(extract_all_text(item))
    return text_list

# Use the recursive function in our loop
post_dfs = {}
for item in json_data['lista']:
    post_id = item['post']['id']
    all_text = extract_all_text(item['post'])  # Grab all text under this post
    post_dfs[post_id] = pd.DataFrame({'Text': all_text})

Step 4: Access Your DataFrames

You can now access each post’s DataFrame using its ID:

# Example: Retrieve the DataFrame for Post ID 'post_456'
print(post_dfs['post_456'])

Quick Notes

  • Adjust the path to Post ID and text in the code to match your actual JSON schema (the examples are generic since I don’t have your exact data structure).
  • If some posts have no text, their DataFrame will be empty—you can add checks to skip those if needed.
  • For very large JSON files, consider using ijson to stream the data instead of loading it all into memory at once.

内容的提问来源于stack exchange,提问作者penestia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:15:34