使用Python+YouTube API v3获取YouTube视频全部评论及回复问题求解
YouTube视频评论及全量回复爬取问题
我正在练习获取YouTube视频的所有评论及对应回复,后续计划基于该代码扩展实现全频道的情感分析和社交网络分析。目前我写的代码没法拉取全量回复,甚至不确定能不能拿到全部主评论,类似问题在Stack Overflow已有讨论但没有给出可行解决方案,代码如下:
import os import googleapiclient.discovery import pandas as pd def main(): # 本地运行时关闭OAuthlib的HTTPS校验 # *生产环境千万不要开启该配置* os.environ["OAUTHLIB_INSECURE_TRANSPORT"] = "1" api_service_name = "youtube" api_version = "v3" DEVELOPER_KEY = "yourapikey" # <--- 在此处填入你的API密钥 youtube = googleapiclient.discovery.build( api_service_name, api_version, developerKey = DEVELOPER_KEY) request = youtube.commentThreads().list( part="snippet, replies", order="time", maxResults=100, textFormat="plainText", videoId="aCjyqziVSMA" ) response = request.execute() full = pd.json_normalize(response, record_path=['items']) while response: if 'nextPageToken' in response: response = youtube.commentThreads().list( part="snippet", maxResults=100, textFormat='plainText', order='time', videoId='aCjyqziVSMA', pageToken=response['nextPageToken'] ).execute() df2 = pd.json_normalize(response, record_path=['items']) full = full.append(df2) else: break return full
运行上述函数后,我用以下代码拆分评论回复,仅选取首个有回复的评论做测试:
df2 = test[test['snippet.totalReplyCount']>0].reset_index(drop=True) pd.json_normalize(df2['replies.comments'][0])
结果显示这条总共有166条回复的评论仅拉取到4条回复。请问我的代码存在什么问题?这是YouTube API本身的调用限制吗?
编辑1
我查阅YouTube官方文档后了解到:若要获取某条顶级评论的全部回复,需要调用comments.list方法并使用parentId请求参数指定要获取回复的评论ID。但我目前对API调用还不够熟悉,不知道如何实现该逻辑,如果有可运行的实现方案我也会采纳。
编辑2(2022-05-23)
以下是我自行摸索出的解决方案:该函数会调用Google API拉取每条评论的所有回复并存入DataFrame,你可以将该数据表与原有评论数据表关联,即可获取全部评论与回复(前提是API配额未耗尽),代码如下:
import os import googleapiclient.discovery import pandas as pd # 回复拉取函数 def repliesto(parentId): # 本地运行时关闭OAuthlib的HTTPS校验 # *生产环境千万不要开启该配置* os.environ["OAUTHLIB_INSECURE_TRANSPORT"] = "1" api_service_name = "youtube" api_version = "v3" DEVELOPER_KEY = DevKey # 你的开发者密钥 youtube = googleapiclient.discovery.build( api_service_name, api_version, developerKey = DEVELOPER_KEY) request = youtube.comments().list( part="snippet", maxResults=100, parentId=parentId, textFormat="plainText" ) response = request.execute() replies = pd.json_normalize(response, record_path=['items']) while response: if 'nextPageToken' in response: response = youtube.comments().list( part="snippet", maxResults=100, parentId=parentId, textFormat="plainText", pageToken=response['nextPageToken'] ).execute() df2 = pd.json_normalize(response, record_path=['items']) replies = pd.concat([replies, df2], sort=False) else: break return replies # 从主评论数据表中提取所有有回复的顶级评论ID,full变量是存储所有主评论的DataFrame replyto = [] for reply in full[(full['snippet.totalReplyCount']>0)]['snippet.topLevelComment.id']: replyto.append(reply) # 新建空DataFrame存储所有回复,遍历replyto列表中的每个ID调用上面定义的拉取函数 replies = pd.DataFrame() for reply in replyto: df = repliesto(reply) replies = pd.concat([replies, df], ignore_index=True)
内容的提问来源于stack exchange,提问作者DataStraine
相关产品推荐
相关产品推荐

