Python如何从YouTube Data API返回的嵌套字典中提取频道评论数据
YouTube Data API嵌套结构评论文本提取问题及解决方案
问题描述
开展研究项目需要获取特定YouTube频道的评论文本数据,已通过Google YouTube Data API完成数据抓取,但返回的数据为复杂的嵌套字典与列表组合结构,暂无法顺利拆解提取目标字段。
目标字段路径:评论的Text Display与Text Original字段存储在Snippet字典中,该字典隶属于Top Level Comments字典,而Top Level Comments是items字典下的列表元素。
现有代码
from googleapiclient.discovery import build api_key = '_______________________________' youtube = build('youtube', 'v3', developerKey = api_key) # 频道ID可通过第三方工具查询获取 request = youtube.commentThreads().list( part = 'snippet', allThreadsRelatedToChannelId = 'UC_zxivooFdvF4uuBosUnJxQ' ) response3 = request.execute() # 数据结构探索代码已省略 # 按指定键筛选字典 includedKeys = ['items'] dataDic = {k:v for k, v in response3.items() if k in includedKeys}
报错信息
尝试提取Snippet列表时触发以下报错:
dataDic2 = {x['snippet'] for x in dataDic} # TypeError: string indices must be integers dataDic2 = [{'snippet': d['snippet']} for d in dataDic] # TypeError: string indices must be integers dataDic2 = [topLevelComment['snippet'] for topLevelComment in dataDic['topLevelComment']['snippet']] # KeyError: 'topLevelComment' import ast result = ast.literal_eval('[snippet]') assert type(result) is list # ValueError: malformed node or string: <_ast.Name object at 0x0000010F6D7B9A08>
参考资源
- 数据结构示意图:

- 样本数据下载地址:sample data
错误原因
核心错误有两点:
dataDic是仅包含items键的字典,直接遍历dataDic时实际遍历的是字典的键(字符串类型的'items'),对字符串使用字典索引就会触发TypeError: string indices must be integers报错- 嵌套层级取值路径错误,
topLevelComment位于items列表每个元素的snippet字段下,跳过中间层级直接取值才会触发KeyError
正确提取代码
易读扩展版(方便后续新增字段)
# 直接从接口返回结果取items列表,无需单独做字典筛选 comment_threads = response3['items'] target_data = [] for thread in comment_threads: # 按层级取值:单条评论线程 → 线程snippet → topLevelComment → 评论内容snippet comment_snippet = thread['snippet']['topLevelComment']['snippet'] # 按需提取目标字段,结构示意图红圈中的其他字段可直接按key添加 target_data.append({ "textDisplay": comment_snippet['textDisplay'], "textOriginal": comment_snippet['textOriginal'] # 示例补充其他红圈字段: # "authorName": comment_snippet['authorDisplayName'], # "publishTime": comment_snippet['publishedAt'], # "likeCount": comment_snippet['likeCount'] })
简写版(列表推导式)
target_data = [ { "textDisplay": thread['snippet']['topLevelComment']['snippet']['textDisplay'], "textOriginal": thread['snippet']['topLevelComment']['snippet']['textOriginal'] } for thread in response3['items'] ]
结果说明
运行完成后target_data就是所有目标字段组成的列表,可直接导出为csv、json等格式用于后续研究分析。
内容的提问来源于stack exchange,提问作者Simone
相关产品推荐
相关产品推荐

