You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

新手求助:如何用BeautifulSoup爬取Xbox游戏的标题、图片等信息?

Xbox官网游戏页面爬取优化方案

问题分析

你当前的代码仅获取了页面中第一个<li>元素,无法批量提取所有游戏信息。需要先定位到所有游戏项的容器,再逐个提取标题、图片、类型等字段。

修改后的代码

import requests
from bs4 import BeautifulSoup

url = "https://www.xbox.com/en-US/browse/games"
# 添加请求头模拟浏览器,避免被反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}
response = requests.get(url, headers=headers)
response.encoding = response.apparent_encoding  # 确保编码正确,避免乱码
soup = BeautifulSoup(response.text, 'lxml')

# 定位所有游戏项的li元素(匹配页面HTML结构中的游戏容器类)
game_items = soup.find_all('li', class_='m-channel-placement-item')

# 遍历每个游戏项提取信息
game_list = []
for item in game_items:
    game_info = {}
    # 提取游戏标题
    title_tag = item.find('h3', class_='c-game-card__title')
    game_info['title'] = title_tag.get_text(strip=True) if title_tag else "无标题"
    
    # 提取游戏图片URL
    img_tag = item.find('img', class_='c-game-card__image')
    game_info['image_url'] = img_tag['src'] if img_tag else "无图片链接"
    
    # 提取游戏类型/副标题
    type_tag = item.find('p', class_='c-game-card__subtitle')
    game_info['genre'] = type_tag.get_text(strip=True) if type_tag else "无类型信息"
    
    game_list.append(game_info)

# 打印提取的结果
for game in game_list:
    print(f"标题:{game['title']}")
    print(f"图片链接:{game['image_url']}")
    print(f"类型:{game['genre']}")
    print("-" * 50)

关键说明

  • 请求头设置:添加User-Agent模拟浏览器请求,避免被网站反爬机制拦截。
  • 批量定位元素:用find_all替代find,一次性获取所有游戏项的容器元素。
  • 容错处理:每个字段提取时先判断标签是否存在,避免因部分游戏信息缺失导致代码报错。
  • 动态加载问题:如果上述代码仅提取到少量游戏,说明页面内容是动态加载的,需要使用selenium等工具模拟浏览器滚动加载更多内容。

内容的提问来源于stack exchange,提问作者ihatecoding

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 10:37:28