无法找到正确的BeautifulSoup类与ID组合,爬取YouTube游戏页面遇阻
问题:无法提取YouTube Gaming页面的游戏名称及实时观看人数
我的代码:
from bs4 import BeautifulSoup import requests URL = 'https://www.youtube.com/gaming/games' response = requests.get(URL).text soup = BeautifulSoup(response, 'html.parser') elem = soup.find_all('a', class_ = 'yt-simple-endpoint focus-on-expand style-scope ytd-game-details-renderer') print(elem)
问题描述:
我想提取YouTube Gaming游戏页面里的所有游戏名称和实时观看人数,但一直找不到正确的标签、class或ID组合。
已尝试的find_all参数:
('a', class_ = 'yt-simple-endpoint focus-on-expand style-scope ytd-game-details-renderer') ('game', class_ = 'style-scope ytd-game-card-renderer') (class_ = 'style-scope ytd-grid-renderer') (id = 'items')
还试过这些参数的多种变体。只用find_all('div')会得到大量无关数据;我觉得id='items'是正确方向,但加上其他参数后都返回空列表[];搜索div结果里的子元素也拿不到有效内容。用find(id='items')返回None,定位id='live-viewers-count'的元素同样返回[]。
页面元素截图:
解决方法
YouTube的页面内容大量依赖JavaScript动态加载,直接用requests.get()只能获取初始静态HTML,动态渲染的游戏数据并不在其中,这是你找不到目标元素的核心原因。以下是两种可行方案:
方案一:使用Selenium模拟浏览器加载
Selenium会启动真实浏览器,等待页面动态加载完成后再获取完整内容,能拿到所有渲染后的元素。
示例代码:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup URL = 'https://www.youtube.com/gaming/games' # 初始化Chrome浏览器(需提前下载对应版本的chromedriver) driver = webdriver.Chrome() driver.get(URL) # 等待游戏卡片元素加载完成,超时时间10秒 wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_element_located((By.TAG_NAME, "ytd-game-card-renderer"))) # 获取完整页面源码 page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') # 提取游戏名称和实时观看人数 game_cards = soup.find_all('ytd-game-card-renderer') for card in game_cards: game_name = card.find('a', class_='yt-simple-endpoint focus-on-expand style-scope ytd-game-details-renderer').text.strip() live_viewers = card.find('span', id='live-viewers-count').text.strip() if card.find('span', id='live-viewers-count') else '无实时数据' print(f"游戏名称:{game_name},实时观看人数:{live_viewers}") # 关闭浏览器 driver.quit()
方案二:调用YouTube Data API
如果不想用浏览器模拟,可直接调用官方API获取数据,这是最稳定的方式,但需要先在Google Cloud平台创建项目,申请API密钥并启用YouTube Data API。
示例代码(需先安装google-api-python-client):
from googleapiclient.discovery import build # 替换为你的API密钥 API_KEY = '你的API密钥' youtube = build('youtube', 'v3', developerKey=API_KEY) # 请求游戏列表数据,可修改regionCode指定地区 request = youtube.games().list( part='snippet,contentDetails', regionCode='US' ) response = request.execute() for item in response['items']: game_name = item['snippet']['title'] print(f"游戏名称:{game_name}")
注意:使用Selenium时,需保证chromedriver版本与Chrome浏览器版本匹配;使用API时,需注意谷歌的API调用配额限制。
内容的提问来源于stack exchange,提问作者user21090678
相关产品推荐
相关产品推荐

