无法获取YouTube视频全部评论:代码卡在99%处停止运行
问题
我用了Stack Overflow上的《How to extract all YouTube comments using YouTube API? (Python)》代码,现在碰到个情况:评论收集到99%就停了,程序既不崩溃也没错误输出,已经确认不是API配额限制的问题。
原代码
def GetVideoComments(youtube, video_id, datetime_to_check, comments=[], token_list = [], token="") : # Stores the total reply count a comment has totalReplyCount = 0 # Replies of the comments replies = [] try : request = youtube.commentThreads().list(part="snippet", videoId=video_id, maxResults=100, pageToken=token) response = request.execute() except HttpError as error : print("Error", error) progress_bar.close() return comments progress_bar.update(len(response["items"])) for item in response["items"] : comment = item["snippet"]["topLevelComment"] timestamp = comment["snippet"]["publishedAt"] # Skip this comment if it is past the time period we want if (timestamp > datetime_to_check) : continue text = comment["snippet"]["textDisplay"] comments.append({"timestamp" : timestamp, "comment" : text, "sentiment" : {}, "meaningful words" : ""}) # Get the total reply count: totalReplyCount = item["snippet"]["totalReplyCount"] # Check if the total reply count is greater than zero, # if so,call GetVideoReplies # and extend the "comments" returned list. if (totalReplyCount > 0) : comments += GetVideoReplies(comment["id"], datetime_to_check, replies, token_list, None) # replies must be cleared as if GetVideoComments is called recursively, # replies will still contain its current elements replies = [] if "nextPageToken" in response and response["nextPageToken"] not in token_list : token_list.append(response["nextPageToken"]) return GetVideoComments(youtube, video_id, datetime_to_check, comments, token_list, response["nextPageToken"]) else : progress_bar.close() return comments def GetVideoReplies(comment_ID, datetime_to_check, replies, token_list, token) : try : request = youtube.comments().list(part="snippet", maxResults=100, parentId=comment_ID, pageToken=token) response = request.execute() except HttpError as error : print("Error", error) progress_bar.close() return replies progress_bar.update(len(response["items"])) for item in response["items"] : # Append the reply's text to replies if (item["snippet"]["publishedAt"] > datetime_to_check) : continue replies.append({"timestamp" : item["snippet"]["publishedAt"], "comment" : item["snippet"]["textDisplay"], "sentiment" : {}, "meaningful words" : ""}) if "nextPageToken" in response and response["nextPageToken"] not in token_list : token_list.append(response["nextPageToken"]) return GetVideoReplies(comment_ID, datetime_to_check, replies, token_list, response["nextPageToken"]) else: return replies
问题排查及修复方案
1. 主评论与回复的分页Token列表共用(核心问题)
原代码里GetVideoComments和GetVideoReplies共用同一个token_list,但两者的分页Token是完全独立的,共用会导致:当回复的Token被加入列表后,主评论如果碰到相同的Token(概率低但存在)会直接停止拉取,或者反过来。
修复:给回复函数单独分配Token列表,不要和主函数共享:
- 调用
GetVideoReplies时,传空列表:comments += GetVideoReplies(comment["id"], datetime_to_check, replies, [], None) - 修改
GetVideoReplies的默认参数,让它自己维护Token列表:
def GetVideoReplies(comment_ID, datetime_to_check, replies=[], token_list=[], token=None) : # 函数内容不变
2. 递归调用的潜在栈溢出风险
Python默认递归深度有限(默认约1000层),如果评论页数极多,递归会触发隐式问题(虽然你没报错,但这是不稳定因素)。换成迭代写法更可靠:
优化后的迭代版代码
def GetVideoComments(youtube, video_id, datetime_to_check): comments = [] token = "" token_list = [] while True: try: request = youtube.commentThreads().list( part="snippet", videoId=video_id, maxResults=100, pageToken=token ) response = request.execute() except HttpError as error: print("请求错误:", error) progress_bar.close() return comments progress_bar.update(len(response["items"])) for item in response["items"]: comment = item["snippet"]["topLevelComment"] timestamp = comment["snippet"]["publishedAt"] if timestamp > datetime_to_check: continue text = comment["snippet"]["textDisplay"] comments.append({ "timestamp": timestamp, "comment": text, "sentiment": {}, "meaningful words": "" }) totalReplyCount = item["snippet"]["totalReplyCount"] if totalReplyCount > 0: # 调用回复函数,传独立的token列表 comments += GetVideoReplies(youtube, comment["id"], datetime_to_check) if "nextPageToken" in response and response["nextPageToken"] not in token_list: token_list.append(response["nextPageToken"]) token = response["nextPageToken"] else: break progress_bar.close() return comments def GetVideoReplies(youtube, comment_ID, datetime_to_check): replies = [] token = "" token_list = [] while True: try: request = youtube.comments().list( part="snippet", maxResults=100, parentId=comment_ID, pageToken=token ) response = request.execute() except HttpError as error: print("请求错误:", error) progress_bar.close() return replies progress_bar.update(len(response["items"])) for item in response["items"]: if item["snippet"]["publishedAt"] > datetime_to_check: continue replies.append({ "timestamp": item["snippet"]["publishedAt"], "comment": item["snippet"]["textDisplay"], "sentiment": {}, "meaningful words": "" }) if "nextPageToken" in response and response["nextPageToken"] not in token_list: token_list.append(response["nextPageToken"]) token = response["nextPageToken"] else: break return replies
注:这里把youtube参数显式传入GetVideoReplies,避免依赖全局变量,代码更健壮。
3. 进度条显示异常排查
如果进度条的总长度是按预估评论数设置的,但实际拉取时因为datetime_to_check过滤了大量评论,会导致进度停在99%。可以加日志排查:
- 在获取每页评论后,打印当前页的评论数和是否有下一页:
print(f"当前页获取{len(response['items'])}条评论,是否有下一页:{'nextPageToken' in response}") - 在过滤评论时打印日志:
if timestamp > datetime_to_check: print(f"跳过评论:{timestamp} 晚于指定时间{datetime_to_check}") continue
这样能快速确认是不是最后一页的评论都被过滤了,导致进度没到100%。
内容的提问来源于stack exchange,提问作者someNoob
相关产品推荐
相关产品推荐

