You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法找到正确的BeautifulSoup类与ID组合,爬取YouTube游戏页面遇阻

问题:无法提取YouTube Gaming页面的游戏名称及实时观看人数

我的代码:

from bs4 import BeautifulSoup
import requests

URL = 'https://www.youtube.com/gaming/games'

response = requests.get(URL).text
soup = BeautifulSoup(response, 'html.parser')

elem = soup.find_all('a', class_ = 'yt-simple-endpoint focus-on-expand style-scope ytd-game-details-renderer')

print(elem)

问题描述:

我想提取YouTube Gaming游戏页面里的所有游戏名称和实时观看人数,但一直找不到正确的标签、class或ID组合。

已尝试的find_all参数:

('a', class_ = 'yt-simple-endpoint focus-on-expand style-scope ytd-game-details-renderer')
('game', class_ = 'style-scope ytd-game-card-renderer')
(class_ = 'style-scope ytd-grid-renderer')
(id = 'items')

还试过这些参数的多种变体。只用find_all('div')会得到大量无关数据;我觉得id='items'是正确方向,但加上其他参数后都返回空列表[];搜索div结果里的子元素也拿不到有效内容。用find(id='items')返回None,定位id='live-viewers-count'的元素同样返回[]。

页面元素截图:
页面元素截图


解决方法

YouTube的页面内容大量依赖JavaScript动态加载,直接用requests.get()只能获取初始静态HTML,动态渲染的游戏数据并不在其中,这是你找不到目标元素的核心原因。以下是两种可行方案:

方案一:使用Selenium模拟浏览器加载

Selenium会启动真实浏览器,等待页面动态加载完成后再获取完整内容,能拿到所有渲染后的元素。

示例代码:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

URL = 'https://www.youtube.com/gaming/games'

# 初始化Chrome浏览器(需提前下载对应版本的chromedriver)
driver = webdriver.Chrome()
driver.get(URL)

# 等待游戏卡片元素加载完成,超时时间10秒
wait = WebDriverWait(driver, 10)
wait.until(EC.presence_of_element_located((By.TAG_NAME, "ytd-game-card-renderer")))

# 获取完整页面源码
page_source = driver.page_source
soup = BeautifulSoup(page_source, 'html.parser')

# 提取游戏名称和实时观看人数
game_cards = soup.find_all('ytd-game-card-renderer')
for card in game_cards:
    game_name = card.find('a', class_='yt-simple-endpoint focus-on-expand style-scope ytd-game-details-renderer').text.strip()
    live_viewers = card.find('span', id='live-viewers-count').text.strip() if card.find('span', id='live-viewers-count') else '无实时数据'
    print(f"游戏名称:{game_name},实时观看人数:{live_viewers}")

# 关闭浏览器
driver.quit()

方案二:调用YouTube Data API

如果不想用浏览器模拟,可直接调用官方API获取数据,这是最稳定的方式,但需要先在Google Cloud平台创建项目,申请API密钥并启用YouTube Data API。

示例代码(需先安装google-api-python-client):

from googleapiclient.discovery import build

# 替换为你的API密钥
API_KEY = '你的API密钥'
youtube = build('youtube', 'v3', developerKey=API_KEY)

# 请求游戏列表数据,可修改regionCode指定地区
request = youtube.games().list(
    part='snippet,contentDetails',
    regionCode='US'
)
response = request.execute()

for item in response['items']:
    game_name = item['snippet']['title']
    print(f"游戏名称:{game_name}")

注意:使用Selenium时,需保证chromedriver版本与Chrome浏览器版本匹配;使用API时,需注意谷歌的API调用配额限制。


内容的提问来源于stack exchange,提问作者user21090678

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 18:39:42