新手求助:如何用BeautifulSoup爬取Xbox游戏的标题、图片等信息?
Xbox官网游戏页面爬取优化方案
问题分析
你当前的代码仅获取了页面中第一个<li>元素,无法批量提取所有游戏信息。需要先定位到所有游戏项的容器,再逐个提取标题、图片、类型等字段。
修改后的代码
import requests from bs4 import BeautifulSoup url = "https://www.xbox.com/en-US/browse/games" # 添加请求头模拟浏览器,避免被反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } response = requests.get(url, headers=headers) response.encoding = response.apparent_encoding # 确保编码正确,避免乱码 soup = BeautifulSoup(response.text, 'lxml') # 定位所有游戏项的li元素(匹配页面HTML结构中的游戏容器类) game_items = soup.find_all('li', class_='m-channel-placement-item') # 遍历每个游戏项提取信息 game_list = [] for item in game_items: game_info = {} # 提取游戏标题 title_tag = item.find('h3', class_='c-game-card__title') game_info['title'] = title_tag.get_text(strip=True) if title_tag else "无标题" # 提取游戏图片URL img_tag = item.find('img', class_='c-game-card__image') game_info['image_url'] = img_tag['src'] if img_tag else "无图片链接" # 提取游戏类型/副标题 type_tag = item.find('p', class_='c-game-card__subtitle') game_info['genre'] = type_tag.get_text(strip=True) if type_tag else "无类型信息" game_list.append(game_info) # 打印提取的结果 for game in game_list: print(f"标题:{game['title']}") print(f"图片链接:{game['image_url']}") print(f"类型:{game['genre']}") print("-" * 50)
关键说明
- 请求头设置:添加
User-Agent模拟浏览器请求,避免被网站反爬机制拦截。 - 批量定位元素:用
find_all替代find,一次性获取所有游戏项的容器元素。 - 容错处理:每个字段提取时先判断标签是否存在,避免因部分游戏信息缺失导致代码报错。
- 动态加载问题:如果上述代码仅提取到少量游戏,说明页面内容是动态加载的,需要使用
selenium等工具模拟浏览器滚动加载更多内容。
内容的提问来源于stack exchange,提问作者ihatecoding
相关产品推荐
相关产品推荐

