You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法获取YouTube视频全部评论:代码卡在99%处停止运行

问题

我用了Stack Overflow上的《How to extract all YouTube comments using YouTube API? (Python)》代码,现在碰到个情况:评论收集到99%就停了,程序既不崩溃也没错误输出,已经确认不是API配额限制的问题。

原代码

def GetVideoComments(youtube, video_id, datetime_to_check, comments=[], token_list = [], token="") :

    # Stores the total reply count a comment has
    totalReplyCount = 0
    
    # Replies of the comments
    replies = []
    
    try :
        request = youtube.commentThreads().list(part="snippet", 
                                                videoId=video_id, 
                                                maxResults=100,
                                                pageToken=token)
        response = request.execute()
    except HttpError as error :  
        print("Error", error)
        progress_bar.close()
        return comments
    
    progress_bar.update(len(response["items"]))
    
    for item in response["items"] :
        
        comment = item["snippet"]["topLevelComment"]
        timestamp = comment["snippet"]["publishedAt"]
        
        # Skip this comment if it is past the time period we want
        if (timestamp > datetime_to_check) :
            continue
            
        text = comment["snippet"]["textDisplay"]
        comments.append({"timestamp" : timestamp,
                         "comment" : text,
                         "sentiment" : {},
                         "meaningful words" : ""})
        
        # Get the total reply count: 
        totalReplyCount = item["snippet"]["totalReplyCount"]

        # Check if the total reply count is greater than zero, 
        # if so,call GetVideoReplies
        # and extend the "comments" returned list.
        if (totalReplyCount > 0) : 
            comments += GetVideoReplies(comment["id"], datetime_to_check, replies, token_list, None)

        # replies must be cleared as if GetVideoComments is called recursively,
        # replies will still contain its current elements
        replies = []

    if "nextPageToken" in response and response["nextPageToken"] not in token_list : 
        token_list.append(response["nextPageToken"])
        return GetVideoComments(youtube, video_id, datetime_to_check, comments, token_list, response["nextPageToken"])
    else :
        progress_bar.close()
        return comments


def GetVideoReplies(comment_ID, datetime_to_check, replies, token_list, token) : 
    
    try :
        request = youtube.comments().list(part="snippet", 
                                          maxResults=100, 
                                          parentId=comment_ID,
                                          pageToken=token)
        response = request.execute()
    except HttpError as error :  
        print("Error", error)
        progress_bar.close()
        return replies

    progress_bar.update(len(response["items"]))
    
    for item in response["items"] :
        
        # Append the reply's text to replies
        if (item["snippet"]["publishedAt"] > datetime_to_check) :
            continue
        replies.append({"timestamp" : item["snippet"]["publishedAt"],
                         "comment" : item["snippet"]["textDisplay"],
                         "sentiment" : {},
                         "meaningful words" : ""})

    if "nextPageToken" in response and response["nextPageToken"] not in token_list : 
        token_list.append(response["nextPageToken"])
        return GetVideoReplies(comment_ID, datetime_to_check, replies, token_list, response["nextPageToken"])
    else:
        return replies

问题排查及修复方案

1. 主评论与回复的分页Token列表共用(核心问题)

原代码里GetVideoComments和GetVideoReplies共用同一个token_list,但两者的分页Token是完全独立的,共用会导致:当回复的Token被加入列表后,主评论如果碰到相同的Token(概率低但存在)会直接停止拉取,或者反过来。

修复:给回复函数单独分配Token列表,不要和主函数共享:

  • 调用GetVideoReplies时,传空列表:comments += GetVideoReplies(comment["id"], datetime_to_check, replies, [], None)
  • 修改GetVideoReplies的默认参数,让它自己维护Token列表:
def GetVideoReplies(comment_ID, datetime_to_check, replies=[], token_list=[], token=None) : 
    # 函数内容不变

2. 递归调用的潜在栈溢出风险

Python默认递归深度有限(默认约1000层),如果评论页数极多,递归会触发隐式问题(虽然你没报错,但这是不稳定因素)。换成迭代写法更可靠:

优化后的迭代版代码

def GetVideoComments(youtube, video_id, datetime_to_check):
    comments = []
    token = ""
    token_list = []
    
    while True:
        try:
            request = youtube.commentThreads().list(
                part="snippet",
                videoId=video_id,
                maxResults=100,
                pageToken=token
            )
            response = request.execute()
        except HttpError as error:
            print("请求错误:", error)
            progress_bar.close()
            return comments
        
        progress_bar.update(len(response["items"]))
        
        for item in response["items"]:
            comment = item["snippet"]["topLevelComment"]
            timestamp = comment["snippet"]["publishedAt"]
            
            if timestamp > datetime_to_check:
                continue
                
            text = comment["snippet"]["textDisplay"]
            comments.append({
                "timestamp": timestamp,
                "comment": text,
                "sentiment": {},
                "meaningful words": ""
            })
            
            totalReplyCount = item["snippet"]["totalReplyCount"]
            if totalReplyCount > 0:
                # 调用回复函数,传独立的token列表
                comments += GetVideoReplies(youtube, comment["id"], datetime_to_check)
        
        if "nextPageToken" in response and response["nextPageToken"] not in token_list:
            token_list.append(response["nextPageToken"])
            token = response["nextPageToken"]
        else:
            break
    
    progress_bar.close()
    return comments


def GetVideoReplies(youtube, comment_ID, datetime_to_check):
    replies = []
    token = ""
    token_list = []
    
    while True:
        try:
            request = youtube.comments().list(
                part="snippet",
                maxResults=100,
                parentId=comment_ID,
                pageToken=token
            )
            response = request.execute()
        except HttpError as error:
            print("请求错误:", error)
            progress_bar.close()
            return replies
        
        progress_bar.update(len(response["items"]))
        
        for item in response["items"]:
            if item["snippet"]["publishedAt"] > datetime_to_check:
                continue
            replies.append({
                "timestamp": item["snippet"]["publishedAt"],
                "comment": item["snippet"]["textDisplay"],
                "sentiment": {},
                "meaningful words": ""
            })
        
        if "nextPageToken" in response and response["nextPageToken"] not in token_list:
            token_list.append(response["nextPageToken"])
            token = response["nextPageToken"]
        else:
            break
    
    return replies

注:这里把youtube参数显式传入GetVideoReplies,避免依赖全局变量,代码更健壮。

3. 进度条显示异常排查

如果进度条的总长度是按预估评论数设置的,但实际拉取时因为datetime_to_check过滤了大量评论,会导致进度停在99%。可以加日志排查:

  • 在获取每页评论后,打印当前页的评论数和是否有下一页:
    print(f"当前页获取{len(response['items'])}条评论,是否有下一页:{'nextPageToken' in response}")
    
  • 在过滤评论时打印日志:
    if timestamp > datetime_to_check:
        print(f"跳过评论:{timestamp} 晚于指定时间{datetime_to_check}")
        continue
    

这样能快速确认是不是最后一页的评论都被过滤了,导致进度没到100%。


内容的提问来源于stack exchange,提问作者someNoob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 06:37:04