You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas处理嵌套JSON:提取行内字典'word'值拼接完整语句

问题背景

现有结构如下的JSON文件:

{
    "result":{
        "segments":[
            {
                "speaker":"#Tag#",
                "language":"",
                "start":0.10,
                "end":0.20,
                "words":[
                    {
                        "start":0.10,
                        "end":0.20,
                        "word":"(music)",
                        "conf":1.00
                    }
                ]
            },
            {
                "speaker":"14",
                "language":"EN",
                "start":0.51,
                "end":7.01,
                "words":[
                    {
                        "start":0.51,
                        "end":0.69,
                        "word":"Some",
                        "conf":1.00
                    },
                    {
                        "start":0.69,
                        "end":0.93,
                        "word":"sentence",
                        "conf":1.00
                    }
                ]
            },
            {
                "speaker":"13",
                "language":"EN",
                "start":7.02,
                "end":127.35,
                "words":[
                    {
                        "start":7.02,
                        "end":7.59,
                        "word":"Some",
                        "conf":0.97
                    },
                    {
                        "start":7.59,
                        "end":7.74,
                        "word":"other",
                        "conf":1.00
                    },
                    {
                        "start":126.93,
                        "end":127.35,
                        "word":"sentence.",
                        "conf":1.00
                    }
                ]
            },
            {
                "speaker":"13",
                "language":"EN",
                "start":671.05,
                "end":671.71,
                "words":[
                    {
                        "start":671.05,
                        "end":671.71,
                        "word":"end",
                        "conf":1.00
                     }
                 ]
             }
          ]
       }
}

当前使用pandas的json_normalize方法做层级归一化处理,执行代码如下:

df = pd.read_json(r'/content/drive/MyDrive/json-to-parse.json')
turns = json_normalize(df['result'])
sentence = json_normalize(turns.T[0])

words = json_normalize(sentence['words'])     
words.insert(0, 'speaker', sentence['speaker'])

代码执行后得到元素为字典类型的DataFrame,需要提取每行所有字典中键为word的对应值,按行拼接得到每行对应的完整语句,目前无法成功提取行内所有字典的word键值。

实现方案

原有写法的问题在于直接对sentence['words']调用json_normalize时,会把所有行的词列表打散成宽表结构,丢失词和原语音片段的对应关系,两种可直接落地的方案如下:

方案1:读取阶段直接拼接文本(代码最简洁)

不需要多次调用json_normalize,加载原始JSON后遍历segments时直接拼接每个片段的完整文本:

import pandas as pd
import json

# 加载原始JSON文件
with open('/content/drive/MyDrive/json-to-parse.json', 'r', encoding='utf-8') as f:
    raw_data = json.load(f)

seg_list = []
for seg in raw_data['result']['segments']:
    # 提取当前片段下所有word值,按空格拼接
    full_sentence = ' '.join([word_item['word'] for word_item in seg['words']])
    seg_list.append({
        'speaker': seg['speaker'],
        'language': seg['language'],
        'start': seg['start'],
        'end': seg['end'],
        'content': full_sentence
    })

result_df = pd.DataFrame(seg_list)

方案2:沿用原有处理流程,行级apply拼接

如果需要保留之前的归一化逻辑,直接在sentence DataFrame上对words列做行级处理即可,不需要额外生成words表:

import pandas as pd
from pandas import json_normalize

df = pd.read_json(r'/content/drive/MyDrive/json-to-parse.json')
turns = json_normalize(df['result'])
sentence = json_normalize(turns.T[0])

# 逐行提取words列表中所有word字段值,拼接成完整语句
sentence['full_content'] = sentence['words'].apply(
    lambda word_col: ' '.join([single_word['word'] for single_word in word_col])
)

如果后续需要保留每个词的时间戳、置信度等细粒度信息,再单独展开words列即可,上述两种方法都不会出现词和原语音片段错位的问题。


内容的提问来源于stack exchange,提问作者JBA

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 01:54:22