Pandas处理嵌套JSON:提取行内字典'word'值拼接完整语句
问题背景
现有结构如下的JSON文件:
{ "result":{ "segments":[ { "speaker":"#Tag#", "language":"", "start":0.10, "end":0.20, "words":[ { "start":0.10, "end":0.20, "word":"(music)", "conf":1.00 } ] }, { "speaker":"14", "language":"EN", "start":0.51, "end":7.01, "words":[ { "start":0.51, "end":0.69, "word":"Some", "conf":1.00 }, { "start":0.69, "end":0.93, "word":"sentence", "conf":1.00 } ] }, { "speaker":"13", "language":"EN", "start":7.02, "end":127.35, "words":[ { "start":7.02, "end":7.59, "word":"Some", "conf":0.97 }, { "start":7.59, "end":7.74, "word":"other", "conf":1.00 }, { "start":126.93, "end":127.35, "word":"sentence.", "conf":1.00 } ] }, { "speaker":"13", "language":"EN", "start":671.05, "end":671.71, "words":[ { "start":671.05, "end":671.71, "word":"end", "conf":1.00 } ] } ] } }
当前使用pandas的json_normalize方法做层级归一化处理,执行代码如下:
df = pd.read_json(r'/content/drive/MyDrive/json-to-parse.json') turns = json_normalize(df['result']) sentence = json_normalize(turns.T[0]) words = json_normalize(sentence['words']) words.insert(0, 'speaker', sentence['speaker'])
代码执行后得到元素为字典类型的DataFrame,需要提取每行所有字典中键为word的对应值,按行拼接得到每行对应的完整语句,目前无法成功提取行内所有字典的word键值。
实现方案
原有写法的问题在于直接对sentence['words']调用json_normalize时,会把所有行的词列表打散成宽表结构,丢失词和原语音片段的对应关系,两种可直接落地的方案如下:
方案1:读取阶段直接拼接文本(代码最简洁)
不需要多次调用json_normalize,加载原始JSON后遍历segments时直接拼接每个片段的完整文本:
import pandas as pd import json # 加载原始JSON文件 with open('/content/drive/MyDrive/json-to-parse.json', 'r', encoding='utf-8') as f: raw_data = json.load(f) seg_list = [] for seg in raw_data['result']['segments']: # 提取当前片段下所有word值,按空格拼接 full_sentence = ' '.join([word_item['word'] for word_item in seg['words']]) seg_list.append({ 'speaker': seg['speaker'], 'language': seg['language'], 'start': seg['start'], 'end': seg['end'], 'content': full_sentence }) result_df = pd.DataFrame(seg_list)
方案2:沿用原有处理流程,行级apply拼接
如果需要保留之前的归一化逻辑,直接在sentence DataFrame上对words列做行级处理即可,不需要额外生成words表:
import pandas as pd from pandas import json_normalize df = pd.read_json(r'/content/drive/MyDrive/json-to-parse.json') turns = json_normalize(df['result']) sentence = json_normalize(turns.T[0]) # 逐行提取words列表中所有word字段值,拼接成完整语句 sentence['full_content'] = sentence['words'].apply( lambda word_col: ' '.join([single_word['word'] for single_word in word_col]) )
如果后续需要保留每个词的时间戳、置信度等细粒度信息,再单独展开words列即可,上述两种方法都不会出现词和原语音片段错位的问题。
内容的提问来源于stack exchange,提问作者JBA
相关产品推荐
相关产品推荐

