You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

读取含多层嵌套列表的JSON转DataFrame 拆解scoring嵌套列

嵌套scoring列展开实现方案

你可以根据自己当前的处理阶段选对应的实现方式,两种方法都不需要额外安装第三方依赖。


方法1:已生成半成品DataFrame时直接处理现有scoring列

如果你已经运行现有代码得到了带嵌套scoring列的DataFrame,不用重新读取原始JSON,直接对列做转换拼接即可,代码最简洁:

import pandas as pd

# 将每行的scoring列表转换为 评分类型:评分值 的字典格式,再展开为独立列
scoring_expand = df['scoring'].apply(
    lambda score_list: {score['provider_type']: score['value'] for score in score_list}
).apply(pd.Series)

# 拼接展开后的评分列,删除原嵌套列
df_final = pd.concat([df.drop(columns=['scoring']), scoring_expand], axis=1)

方法2:读取原始JSON时一次性完成全量展开

如果还在数据导入阶段,可以直接在json_normalize步骤处理完两层嵌套,再转成宽表,容错性更高:

import pandas as pd
from pandas import json_normalize

# 第一步:将两层嵌套列表全部打平为长表格式
df_long = pd.concat([
    json_normalize(
        data=entry,
        record_path=['items', 'scoring'], # 指定最深层嵌套列表的路径
        meta=[ # 声明需要保留的非列表字段
            'page', 'page_size', 'total_pages', 'total_results',
            ['items', 'jw_entity_id'], ['items', 'id'], ['items', 'title'], ['items', 'object_type']
        ],
        errors='ignore' # 遇到缺字段的异常数据自动跳过,不中断运行
    ) for entry in my_data
])

# 第二步:将不同评分类型从行转为独立列,得到最终宽表
df_final = df_long.pivot(
    index=[
        'items.jw_entity_id', 'items.id', 'items.title', 'items.object_type',
        'page', 'page_size', 'total_pages', 'total_results'
    ],
    columns='provider_type',
    values='value'
).reset_index().rename_axis(columns=None) # 清理pivot生成的多余索引名

处理完成后,imdb:score、tmdb:score、imdb:votes、tmdb:id都会成为独立列,单条作品占一行,缺失对应评分的位置会自动填充NaN。

内容的提问来源于stack exchange,提问作者interferemadly

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.02 06:21:59