You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何从YouTube Data API返回的嵌套字典中提取频道评论数据

YouTube Data API嵌套结构评论文本提取问题及解决方案

问题描述

开展研究项目需要获取特定YouTube频道的评论文本数据,已通过Google YouTube Data API完成数据抓取,但返回的数据为复杂的嵌套字典与列表组合结构,暂无法顺利拆解提取目标字段。
目标字段路径:评论的Text Display与Text Original字段存储在Snippet字典中,该字典隶属于Top Level Comments字典,而Top Level Comments是items字典下的列表元素。

现有代码

from googleapiclient.discovery import build
api_key = '_______________________________'

youtube = build('youtube', 'v3', developerKey = api_key)

# 频道ID可通过第三方工具查询获取
request = youtube.commentThreads().list(
    part = 'snippet',
    allThreadsRelatedToChannelId = 'UC_zxivooFdvF4uuBosUnJxQ'
    )

response3 = request.execute()

# 数据结构探索代码已省略
# 按指定键筛选字典
includedKeys = ['items']
dataDic = {k:v for k, v in response3.items() if k in includedKeys}

报错信息

尝试提取Snippet列表时触发以下报错:

dataDic2 = {x['snippet'] for x in dataDic} 
# TypeError: string indices must be integers

dataDic2 = [{'snippet': d['snippet']} for d in dataDic] 
# TypeError: string indices must be integers 

dataDic2 = [topLevelComment['snippet'] for topLevelComment in dataDic['topLevelComment']['snippet']] 
# KeyError: 'topLevelComment'

import ast
result = ast.literal_eval('[snippet]') 
assert type(result) is list 
# ValueError: malformed node or string: <_ast.Name object at 0x0000010F6D7B9A08>

参考资源

  • 数据结构示意图:data structure
  • 样本数据下载地址:sample data

错误原因

核心错误有两点:

  1. dataDic是仅包含items键的字典,直接遍历dataDic时实际遍历的是字典的键(字符串类型的'items'),对字符串使用字典索引就会触发TypeError: string indices must be integers报错
  2. 嵌套层级取值路径错误,topLevelComment位于items列表每个元素的snippet字段下,跳过中间层级直接取值才会触发KeyError

正确提取代码

易读扩展版(方便后续新增字段)

# 直接从接口返回结果取items列表,无需单独做字典筛选
comment_threads = response3['items']
target_data = []

for thread in comment_threads:
    # 按层级取值:单条评论线程 → 线程snippet → topLevelComment → 评论内容snippet
    comment_snippet = thread['snippet']['topLevelComment']['snippet']
    # 按需提取目标字段,结构示意图红圈中的其他字段可直接按key添加
    target_data.append({
        "textDisplay": comment_snippet['textDisplay'],
        "textOriginal": comment_snippet['textOriginal']
        # 示例补充其他红圈字段:
        # "authorName": comment_snippet['authorDisplayName'],
        # "publishTime": comment_snippet['publishedAt'],
        # "likeCount": comment_snippet['likeCount']
    })

简写版(列表推导式)

target_data = [
    {
        "textDisplay": thread['snippet']['topLevelComment']['snippet']['textDisplay'],
        "textOriginal": thread['snippet']['topLevelComment']['snippet']['textOriginal']
    }
    for thread in response3['items']
]

结果说明

运行完成后target_data就是所有目标字段组成的列表,可直接导出为csv、json等格式用于后续研究分析。


内容的提问来源于stack exchange,提问作者Simone

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 18:27:03