You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Youtube Data API获取指定时段最高播放量视频结果不准确咨询

问题:Youtube Data API 获取指定时间段高播放量视频结果不准确

我正在探索Youtube Data API,尝试获取某一时间段内发布的最高播放量视频,但得到的结果并不准确。以下是我的代码:

from datetime import datetime
import requests
import json

def get_most_watched_videos(api_key, year_b, month_b, day_b, hour_b, min_b, year_a, month_a, day_a, hour_a, min_a, category_id):
    published_before = datetime(year_b, month_b, day_b, hour_b, min_b).isoformat("T") + "Z"
    published_after = datetime(year_a, month_a, day_a, hour_a, min_a).isoformat("T") + "Z"

    params = {
        'part': 'snippet',
        'maxResults': 10,
        'order': 'viewCount',
        'type': 'video',
        'publishedBefore': published_before,
        'publishedAfter': published_after,
        'region': 'US',
        #'relevanceLanguage': 'en',
        'key': api_key,
    }

    response = requests.get('https://www.googleapis.com/youtube/v3/search', params=params)

    return response.json()


year_b, month_b, day_b, hour_b, min_b = 2023, 6, 25, 0, 0
year_a, month_a, day_a, hour_a, min_a = 2010, 6, 24, 0, 0
videos = get_most_watched_videos(api_key, year_b, month_b, day_b, hour_b, min_b, year_a, month_a, day_a, hour_a, min_a, category_id)
video_ids = [video['id']['videoId'] for video in videos['items']]

# get the view count for each video
view_counts = []
video_list = []
for video_id in video_ids:
    response = requests.get(f'https://www.googleapis.com/youtube/v3/videos?part=statistics&id={video_id}&key={api_key}')
    video = json.loads(response.text)
    video_list.append(video)
    view_count = int(video['items'][0]['statistics']['viewCount'])
    view_counts.append((video_id, view_count))

# sort the videos by view count
view_counts.sort(key=lambda x: x[1], reverse=True)

# print the videos and their view counts
for video_id, view_count in view_counts:
    print(f'Video ID: {video_id}, View Count: {view_count}')

该函数返回的结果与实际播放量不符,比如《Despacito》的播放量远超过1700万,但结果中最高播放量仅约1700万。我是否错误使用了API,还是API本身存在问题?

我曾搜索相关问题,找到过类似“youtube data api search by viewCount wrong results”的讨论,但其中推荐使用已废弃的v1版本API。此外,当我扩大时间范围时,得到的播放量结果反而更低。


解决方案

核心问题:Search API的order=viewCount并非全局排序

Youtube Data API的search端点使用order=viewCount时,不是对全网符合条件的视频按播放量全局排序,而是基于Youtube的内部相关性算法,结合播放量、地域、内容匹配等因素返回结果。这就是为什么你拿不到真正的Top播放量视频,甚至扩大时间范围结果反而更差——范围越大,相关性筛选的权重占比越高,真正的高播放量视频可能因为“相关性”低被过滤。

正确的做法:结合Search和Videos端点,分步骤优化

  1. 避免直接用Search端点排序
    Search端点的排序逻辑不是纯粹按播放量,所以不要依赖它返回Top播放量视频。可以先通过Search获取一批候选视频,但要注意:

    • 不要过度依赖order=viewCount,可以先用order=relevance或者order=publishedAt获取更多候选,再自行排序
    • 增大maxResults(最大50),或者通过分页(pageToken)获取更多候选,避免漏过真正的高播放量视频
  2. 直接用Videos端点获取统计数据并排序
    如果你已经有一批候选视频ID,直接调用Videos端点获取statistics部分,然后自己按viewCount排序。但如果没有候选ID,需要先通过Search构建候选池。

  3. 修复你的代码逻辑
    你的代码里有几个小问题:

    • category_id参数定义了但没用到,应该加到Search的params里
    • region参数应该是regionCode(正确的参数名是regionCode,不是region)
    • 时间范围的变量名逻辑没问题,但要确保参数传递对应正确

    优化后的代码示例:

from datetime import datetime
import requests
import json

def get_video_candidates(api_key, published_after, published_before, category_id, max_results=50):
    params = {
        'part': 'snippet',
        'maxResults': max_results,
        'type': 'video',
        'publishedBefore': published_before,
        'publishedAfter': published_after,
        'regionCode': 'US',
        'videoCategoryId': category_id,
        'key': api_key,
    }
    
    all_items = []
    next_page_token = None
    
    # 分页获取更多候选视频,避免漏过优质内容
    while True:
        if next_page_token:
            params['pageToken'] = next_page_token
        response = requests.get('https://www.googleapis.com/youtube/v3/search', params=params)
        data = response.json()
        all_items.extend(data.get('items', []))
        
        next_page_token = data.get('nextPageToken')
        # 限制最多获取200个,避免API配额耗尽
        if not next_page_token or len(all_items) >= 200:
            break
    
    return [item['id']['videoId'] for item in all_items]

def get_video_statistics(api_key, video_ids):
    # 批量查询视频统计,每次最多50个ID,减少请求次数
    stats = []
    for i in range(0, len(video_ids), 50):
        batch_ids = ','.join(video_ids[i:i+50])
        response = requests.get(f'https://www.googleapis.com/youtube/v3/videos?part=statistics,snippet&id={batch_ids}&key={api_key}')
        data = response.json()
        for item in data.get('items', []):
            stats.append({
                'videoId': item['id'],
                'title': item['snippet']['title'],
                'viewCount': int(item['statistics']['viewCount'])
            })
    return stats

# 配置参数
api_key = '你的API密钥'
category_id = '10'  # 示例:音乐分类ID
published_after = datetime(2010, 6, 24, 0, 0).isoformat("T") + "Z"
published_before = datetime(2023, 6, 25, 0, 0).isoformat("T") + "Z"

# 获取候选视频ID
video_ids = get_video_candidates(api_key, published_after, published_before, category_id)
# 获取统计数据
video_stats = get_video_statistics(api_key, video_ids)
# 按播放量降序排序
video_stats.sort(key=lambda x: x['viewCount'], reverse=True)

# 输出Top10结果
for idx, video in enumerate(video_stats[:10]):
    print(f"{idx+1}. {video['title']} (ID: {video['videoId']}) - 播放量: {video['viewCount']:,}")

额外注意事项

  • API配额限制:批量查询可以减少请求次数,避免快速耗尽配额(Search每次请求消耗100单位配额,Videos每次仅消耗1单位)
  • 地域限制:regionCode会影响返回的视频范围,如果需要全球结果,可以去掉该参数
  • 没有完美的全局排序:Youtube没有开放直接获取全网Top播放量视频的API,只能通过构建足够大的候选池来接近真实结果
  • 废弃API不要用:v1版本早已停止服务,使用会导致请求失败

内容的提问来源于stack exchange,提问作者Bernardo Barias

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 18:57:56