将语音识别结果words字段数据转换为DataFrame的Python代码
问题
给定语音识别结果字典result,其segments数组内的每个元素都包含words字段,该字段是一组字典构成的列表,每个字典包含text、start、end等键。需要编写Python代码,提取所有words里的text、start、end数据并转换为Pandas DataFrame。
result的结构示例如下:
result={ "text": " A nation is a community of people formed on the basis of a combination of shared features. A nation is thus the collective identity of a group of people understood as defined by those features.", "segments": [ { "id": 0, "seek": 0, "start": 0.54, "end": 5.08, "text": " A nation is a community of people formed on the basis of a combination of shared features.", "tokens": [50364,13858,], "temperature": 0.0, "avg_logprob": -0.3795953684839709, "compression_ratio": 1.9613899613899615, "no_speech_prob": 0.0013710602652281523, "confidence": 0.607, "words": [ {"text": "a", "start": 0.54, "end": 1.3, "confidence": 0.613}, {"text": "nation", "start": 1.3, "end": 1.98, "confidence": 0.585}, {"text": "is", "start": 1.98, "end": 2.54, "confidence": 0.974}, {"text": "community", "start": 2.54, "end": 3.04, "confidence": 0.808}, {"text": "of", "start": 3.04, "end": 3.48, "confidence": 0.807}, {"text": "people", "start": 3.48, "end": 3.86, "confidence": 0.764}, {"text": "form", "start": 3.86, "end": 4.24, "confidence": 0.477}, {"text": "on.", "start": 4.24, "end": 5.08, "confidence": 0.29} ] }, { "id": 1, "seek": 0, "start": 5.38, "end": 9.72, "text": " the basis of a combination of shared features", "tokens": [50614,4704], "temperature": 0.0, "avg_logprob": -0.3795953684839709, "compression_ratio": 1.9613899613899615, "no_speech_prob": 0.0013710602652281523, "confidence": 0.865, "words": [ {"text": "the", "start": 5.38, "end": 5.8, "confidence": 0.813}, {"text": "basis", "start": 5.8, "end": 6.2, "confidence": 0.886}, {"text": "of", "start": 6.2, "end": 6.54, "confidence": 0.961}, {"text": "a", "start": 6.54, "end": 7.14, "confidence": 0.612}, {"text": "combination", "start": 7.14, "end": 7.8, "confidence": 0.887}, {"text": "of", "start": 7.8, "end": 8.18, "confidence": 0.971}, {"text": "shared", "start": 8.18, "end": 8.5, "confidence": 0.895}, {"text": "features", "start": 8.5, "end": 8.72, "confidence": 0.88} ] }, {"id": 2,...} ], "language": "ko" }
解决方案
通过遍历segments数组,逐个提取每个分段里的words列表,再从每个单词字典中筛选目标字段,最后用Pandas转换为DataFrame即可。
代码实现
import pandas as pd # 初始化空列表存储单词数据 word_data = [] # 遍历所有分段 for segment in result["segments"]: # 遍历当前分段下的每个单词 for word in segment["words"]: # 提取需要的字段并加入列表 word_data.append({ "text": word["text"], "start": word["start"], "end": word["end"] }) # 转换为DataFrame df = pd.DataFrame(word_data) # 可选:打印查看结果 print(df)
代码说明
- 导入
pandas库,这是处理结构化表格数据的核心工具。 - 创建空列表
word_data,用来统一收集所有符合要求的单词信息。 - 嵌套遍历:先遍历所有语音分段,再遍历每个分段下的单词字典。
- 从每个单词字典中提取
text、start、end三个字段,组装成新字典后加入列表。 - 用
pd.DataFrame()将列表转换为结构化的DataFrame,方便后续分析或导出。
扩展优化
如果需要保留confidence等其他字段,只需在添加字典时补充对应键值对即可,示例如下:
word_data.append({ "text": word["text"], "start": word["start"], "end": word["end"], "confidence": word["confidence"] })
内容的提问来源于stack exchange,提问作者user20189397
相关产品推荐
相关产品推荐

