You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python+YouTube API v3获取YouTube视频全部评论及回复问题求解

YouTube视频评论及全量回复爬取问题

我正在练习获取YouTube视频的所有评论及对应回复,后续计划基于该代码扩展实现全频道的情感分析和社交网络分析。目前我写的代码没法拉取全量回复,甚至不确定能不能拿到全部主评论,类似问题在Stack Overflow已有讨论但没有给出可行解决方案,代码如下:

import os
import googleapiclient.discovery
import pandas as pd

def main():
    # 本地运行时关闭OAuthlib的HTTPS校验
    # *生产环境千万不要开启该配置*
    os.environ["OAUTHLIB_INSECURE_TRANSPORT"] = "1"

    api_service_name = "youtube"
    api_version = "v3"
    DEVELOPER_KEY = "yourapikey" # <--- 在此处填入你的API密钥

    youtube = googleapiclient.discovery.build(
        api_service_name, api_version, developerKey = DEVELOPER_KEY)

    request = youtube.commentThreads().list(
        part="snippet, replies",
        order="time",
        maxResults=100,
        textFormat="plainText",
        videoId="aCjyqziVSMA"
    )
    
    response = request.execute()
    full = pd.json_normalize(response, record_path=['items'])
    while response:
        
        if 'nextPageToken' in response:
            response = youtube.commentThreads().list(
                part="snippet",
                maxResults=100,
                textFormat='plainText',
                order='time',
                videoId='aCjyqziVSMA',
                pageToken=response['nextPageToken']
            ).execute()
            
            df2 = pd.json_normalize(response, record_path=['items'])
            full = full.append(df2)
            
        else:
            break
    return full

运行上述函数后,我用以下代码拆分评论回复,仅选取首个有回复的评论做测试:

df2 = test[test['snippet.totalReplyCount']>0].reset_index(drop=True)
pd.json_normalize(df2['replies.comments'][0])

结果显示这条总共有166条回复的评论仅拉取到4条回复。请问我的代码存在什么问题?这是YouTube API本身的调用限制吗?

编辑1

我查阅YouTube官方文档后了解到:若要获取某条顶级评论的全部回复,需要调用comments.list方法并使用parentId请求参数指定要获取回复的评论ID。但我目前对API调用还不够熟悉,不知道如何实现该逻辑,如果有可运行的实现方案我也会采纳。

编辑2(2022-05-23)

以下是我自行摸索出的解决方案:该函数会调用Google API拉取每条评论的所有回复并存入DataFrame,你可以将该数据表与原有评论数据表关联,即可获取全部评论与回复(前提是API配额未耗尽),代码如下:

import os
import googleapiclient.discovery
import pandas as pd

# 回复拉取函数
def repliesto(parentId):
    # 本地运行时关闭OAuthlib的HTTPS校验
    # *生产环境千万不要开启该配置*
    os.environ["OAUTHLIB_INSECURE_TRANSPORT"] = "1"

    api_service_name = "youtube"
    api_version = "v3"
    DEVELOPER_KEY = DevKey # 你的开发者密钥

    youtube = googleapiclient.discovery.build(
        api_service_name, api_version, developerKey = DEVELOPER_KEY)

    request = youtube.comments().list(
        part="snippet",
        maxResults=100,
        parentId=parentId,
        textFormat="plainText"
    )
    response = request.execute()

    replies = pd.json_normalize(response, record_path=['items'])
    while response:

        if 'nextPageToken' in response:
            response = youtube.comments().list(
                part="snippet",
                maxResults=100,
                parentId=parentId,
                textFormat="plainText",
                pageToken=response['nextPageToken']                
            ).execute()

            df2 = pd.json_normalize(response, record_path=['items'])
            replies = pd.concat([replies, df2], sort=False)

        else:
            break
    return replies

# 从主评论数据表中提取所有有回复的顶级评论ID,full变量是存储所有主评论的DataFrame
replyto = []
for reply in full[(full['snippet.totalReplyCount']>0)]['snippet.topLevelComment.id']:
    replyto.append(reply)

# 新建空DataFrame存储所有回复,遍历replyto列表中的每个ID调用上面定义的拉取函数
replies = pd.DataFrame()
for reply in replyto:
    df = repliesto(reply)
    replies = pd.concat([replies, df], ignore_index=True)

内容的提问来源于stack exchange,提问作者DataStraine

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 01:51:03