使用Youtube Data API获取指定时段最高播放量视频结果不准确咨询
我正在探索Youtube Data API,尝试获取某一时间段内发布的最高播放量视频,但得到的结果并不准确。以下是我的代码:
from datetime import datetime import requests import json def get_most_watched_videos(api_key, year_b, month_b, day_b, hour_b, min_b, year_a, month_a, day_a, hour_a, min_a, category_id): published_before = datetime(year_b, month_b, day_b, hour_b, min_b).isoformat("T") + "Z" published_after = datetime(year_a, month_a, day_a, hour_a, min_a).isoformat("T") + "Z" params = { 'part': 'snippet', 'maxResults': 10, 'order': 'viewCount', 'type': 'video', 'publishedBefore': published_before, 'publishedAfter': published_after, 'region': 'US', #'relevanceLanguage': 'en', 'key': api_key, } response = requests.get('https://www.googleapis.com/youtube/v3/search', params=params) return response.json() year_b, month_b, day_b, hour_b, min_b = 2023, 6, 25, 0, 0 year_a, month_a, day_a, hour_a, min_a = 2010, 6, 24, 0, 0 videos = get_most_watched_videos(api_key, year_b, month_b, day_b, hour_b, min_b, year_a, month_a, day_a, hour_a, min_a, category_id) video_ids = [video['id']['videoId'] for video in videos['items']] # get the view count for each video view_counts = [] video_list = [] for video_id in video_ids: response = requests.get(f'https://www.googleapis.com/youtube/v3/videos?part=statistics&id={video_id}&key={api_key}') video = json.loads(response.text) video_list.append(video) view_count = int(video['items'][0]['statistics']['viewCount']) view_counts.append((video_id, view_count)) # sort the videos by view count view_counts.sort(key=lambda x: x[1], reverse=True) # print the videos and their view counts for video_id, view_count in view_counts: print(f'Video ID: {video_id}, View Count: {view_count}')
该函数返回的结果与实际播放量不符,比如《Despacito》的播放量远超过1700万,但结果中最高播放量仅约1700万。我是否错误使用了API,还是API本身存在问题?
我曾搜索相关问题,找到过类似“youtube data api search by viewCount wrong results”的讨论,但其中推荐使用已废弃的v1版本API。此外,当我扩大时间范围时,得到的播放量结果反而更低。
核心问题:Search API的order=viewCount并非全局排序
Youtube Data API的search端点使用order=viewCount时,不是对全网符合条件的视频按播放量全局排序,而是基于Youtube的内部相关性算法,结合播放量、地域、内容匹配等因素返回结果。这就是为什么你拿不到真正的Top播放量视频,甚至扩大时间范围结果反而更差——范围越大,相关性筛选的权重占比越高,真正的高播放量视频可能因为“相关性”低被过滤。
正确的做法:结合Search和Videos端点,分步骤优化
避免直接用Search端点排序
Search端点的排序逻辑不是纯粹按播放量,所以不要依赖它返回Top播放量视频。可以先通过Search获取一批候选视频,但要注意:- 不要过度依赖
order=viewCount,可以先用order=relevance或者order=publishedAt获取更多候选,再自行排序 - 增大
maxResults(最大50),或者通过分页(pageToken)获取更多候选,避免漏过真正的高播放量视频
- 不要过度依赖
直接用Videos端点获取统计数据并排序
如果你已经有一批候选视频ID,直接调用Videos端点获取statistics部分,然后自己按viewCount排序。但如果没有候选ID,需要先通过Search构建候选池。修复你的代码逻辑
你的代码里有几个小问题:category_id参数定义了但没用到,应该加到Search的params里region参数应该是regionCode(正确的参数名是regionCode,不是region)- 时间范围的变量名逻辑没问题,但要确保参数传递对应正确
优化后的代码示例:
from datetime import datetime import requests import json def get_video_candidates(api_key, published_after, published_before, category_id, max_results=50): params = { 'part': 'snippet', 'maxResults': max_results, 'type': 'video', 'publishedBefore': published_before, 'publishedAfter': published_after, 'regionCode': 'US', 'videoCategoryId': category_id, 'key': api_key, } all_items = [] next_page_token = None # 分页获取更多候选视频,避免漏过优质内容 while True: if next_page_token: params['pageToken'] = next_page_token response = requests.get('https://www.googleapis.com/youtube/v3/search', params=params) data = response.json() all_items.extend(data.get('items', [])) next_page_token = data.get('nextPageToken') # 限制最多获取200个,避免API配额耗尽 if not next_page_token or len(all_items) >= 200: break return [item['id']['videoId'] for item in all_items] def get_video_statistics(api_key, video_ids): # 批量查询视频统计,每次最多50个ID,减少请求次数 stats = [] for i in range(0, len(video_ids), 50): batch_ids = ','.join(video_ids[i:i+50]) response = requests.get(f'https://www.googleapis.com/youtube/v3/videos?part=statistics,snippet&id={batch_ids}&key={api_key}') data = response.json() for item in data.get('items', []): stats.append({ 'videoId': item['id'], 'title': item['snippet']['title'], 'viewCount': int(item['statistics']['viewCount']) }) return stats # 配置参数 api_key = '你的API密钥' category_id = '10' # 示例:音乐分类ID published_after = datetime(2010, 6, 24, 0, 0).isoformat("T") + "Z" published_before = datetime(2023, 6, 25, 0, 0).isoformat("T") + "Z" # 获取候选视频ID video_ids = get_video_candidates(api_key, published_after, published_before, category_id) # 获取统计数据 video_stats = get_video_statistics(api_key, video_ids) # 按播放量降序排序 video_stats.sort(key=lambda x: x['viewCount'], reverse=True) # 输出Top10结果 for idx, video in enumerate(video_stats[:10]): print(f"{idx+1}. {video['title']} (ID: {video['videoId']}) - 播放量: {video['viewCount']:,}")
额外注意事项
- API配额限制:批量查询可以减少请求次数,避免快速耗尽配额(Search每次请求消耗100单位配额,Videos每次仅消耗1单位)
- 地域限制:
regionCode会影响返回的视频范围,如果需要全球结果,可以去掉该参数 - 没有完美的全局排序:Youtube没有开放直接获取全网Top播放量视频的API,只能通过构建足够大的候选池来接近真实结果
- 废弃API不要用:v1版本早已停止服务,使用会导致请求失败
内容的提问来源于stack exchange,提问作者Bernardo Barias

