You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python调用YouTube API抓取视频评论及回复出现无限循环问题求助

Python调用YouTube Data API v3抓取视频评论无限循环问题修复

问题根源

你的代码死循环、重复拉取相同评论的核心原因是翻页逻辑漏传关键参数:

  • 检测到响应中的nextPageToken后,你发起的下一页请求没有把这个token作为pageToken传入接口,API每次都会返回第一页数据,循环永远不会触发终止条件。

另外两个会导致代码运行异常/结果不准的问题:

  • 缩进错误:函数定义后的所有代码没有做层级缩进,Python运行时会直接抛出语法错误
  • 回复抓取不全:commentThreads.list接口默认最多随父评论返回5条回复,单条评论回复数超过5条时,剩余回复不会出现在当前响应中,现有代码会漏抓这部分内容

修正后可运行代码

from googleapiclient.discovery import build

# 替换为你自己申请的API密钥
api_key = "YOUR_API_KEY"

def video_comments(video_id):
    all_comments = []
    # 初始化API客户端
    youtube = build('youtube', 'v3', developerKey=api_key)

    # 发起第一页请求
    video_response = youtube.commentThreads().list(
        part='snippet,replies',
        videoId=video_id,
        maxResults=100, # 单页最大拉取数,上限100,可减少请求次数
        order='time' # 排序规则:time按时间倒序,relevance按相关性排序
    ).execute()

    while video_response:
        for item in video_response['items']:
            # 提取顶层评论内容
            top_comment_text = item['snippet']['topLevelComment']['snippet']['textDisplay']
            current_comment = {
                "content": top_comment_text,
                "replies": []
            }
            reply_count = item['snippet']['totalReplyCount']

            if reply_count > 0:
                # 先提取当前响应中附带的回复
                for reply in item['replies']['comments']:
                    reply_text = reply['snippet']['textDisplay']
                    current_comment["replies"].append(reply_text)
                
                # 如果附带回复数小于总回复数,单独拉取全量回复
                if len(item['replies']['comments']) < reply_count:
                    reply_req = youtube.comments().list(
                        part='snippet',
                        parentId=item['id'],
                        maxResults=100
                    )
                    # 回复列表同样需要翻页拉取
                    while reply_req:
                        reply_resp = reply_req.execute()
                        for reply in reply_resp['items']:
                            current_comment["replies"].append(reply['snippet']['textDisplay'])
                        reply_req = youtube.comments().list_next(reply_req, reply_resp)
            
            print(f"评论:{top_comment_text}\n对应回复:{current_comment['replies']}\n")
            all_comments.append(current_comment)

        # 自动处理翻页token,不需要手动传pageToken参数
        video_response = youtube.commentThreads().list_next(video_response, video_response)

    return all_comments

# 调用示例,对应你提供的视频ID
if __name__ == "__main__":
    result = video_comments("kqtD5dpn9C8")
    print(f"抓取完成,共获取{len(result)}条顶层评论")

运行注意事项

  • 提前安装依赖:执行pip install google-api-python-client安装官方API客户端
  • 接口有每日配额限制,默认每天10000单位配额,拉取100条顶层评论/100条回复各消耗1单位,普通视频评论抓取基本不会触发配额上限
  • 若目标视频关闭评论、评论处于审核状态,接口会返回对应错误,可自行添加try/except逻辑做异常处理

内容的提问来源于stack exchange,提问作者user19389509

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 19:39:37