Steam市场爬虫仅抓取第一页问题求助(附代码)
Steam市场爬虫多页抓取异常排查与修复
核心问题定位
你遇到的问题根源是URL中的锚点(#p6_popular_desc这类格式)不会被服务器接收。HTTP请求中,锚点属于浏览器端的定位标记,不会随请求发送到Steam服务器,所以无论你修改#后的内容,服务器始终返回第一页的默认数据。
具体问题与修复方案
分页参数传递错误
Steam市场的分页参数是通过查询字符串(URL中?后的部分)传递的,而非锚点。正确的分页URL格式应为:https://steamcommunity.com/market/search?q=&appid=730&p=6&sort_column=popular&sort_dir=desc其中
p=6表示第6页,sort_column=popular和sort_dir=desc对应你原来锚点里的popular_desc排序规则。headers变量未定义
你的getData函数中使用了headers参数,但代码里没有定义这个变量,会导致请求失败或被Steam反爬机制拦截。需要添加合理的请求头,比如模拟浏览器:headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' }分页函数逻辑错误
getNextPage和isNextPage依赖锚点提取当前页的逻辑完全无效,需要改为从页面的分页按钮中提取下一页参数,或者直接递增分页数字构造正确URL。价格提取函数的bug
lowestPriceExtractor中,find(class_="sale_price")返回单个元素,你却用for循环遍历它,会导致遍历字符而非有效内容,应该直接提取文本:def lowestPriceExtractor(inList): lowestpriceSoup = inList.find(class_="sale_price") if not lowestpriceSoup: return 0.0 lowestpriceval = lowestpriceSoup.get_text(strip=True) lowestpriceret = lowestpriceval.replace("$", '').replace("USD", '').strip() return float(lowestpriceret) if lowestpriceret else 0.0highestPriceExtractor用字符串切片提取数据的方式非常脆弱,容易因HTML结构变化失效,建议改用BeautifulSoup的API提取:def highestPriceExtractor(inList): highestpriceSoup = inList.find(class_="normal_price") if not highestpriceSoup: return 0.0 highestpriceval = highestpriceSoup.get_text(strip=True) highestpriceret = highestpriceval.replace("$", '').replace("USD", '').strip() return float(highestpriceret) if highestpriceret else 0.0
修复后的完整代码示例
from bs4 import BeautifulSoup import numpy as np import requests # 定义请求头,模拟浏览器 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } def highestPriceExtractor(inList): highestpriceSoup = inList.find(class_="normal_price") if not highestpriceSoup: return 0.0 highestpriceval = highestpriceSoup.get_text(strip=True) highestpriceret = highestpriceval.replace("$", '').replace("USD", '').strip() return float(highestpriceret) if highestpriceret else 0.0 def lowestPriceExtractor(inList): lowestpriceSoup = inList.find(class_="sale_price") if not lowestpriceSoup: return 0.0 lowestpriceval = lowestpriceSoup.get_text(strip=True) lowestpriceret = lowestpriceval.replace("$", '').replace("USD", '').strip() return float(lowestpriceret) if lowestpriceret else 0.0 def getNextPage(current_page): # 直接递增页码构造下一页URL,也可从页面分页按钮提取 next_page = current_page + 1 return f"https://steamcommunity.com/market/search?q=&appid=730&p={next_page}&sort_column=popular&sort_dir=desc" def pageNameRetrieve(list_items): nameArr = [] for item in list_items: name = item.get("data-hash-name") if name: nameArr.append(name) return nameArr def pageValueRetrieve(list_items): priceArr = [] for item in list_items: lowestprice = lowestPriceExtractor(item) highestprice = highestPriceExtractor(item) priceArr.append([lowestprice, highestprice]) return np.array(priceArr, dtype=float) def getData(url): try: r = requests.get(url, headers=headers) r.raise_for_status() # 检查请求是否成功 soup = BeautifulSoup(r.text, "html.parser") return soup except requests.exceptions.RequestException as e: print(f"请求失败: {e}") return None # 测试第6页数据 current_page = 6 url = getNextPage(current_page - 1) # 获取第6页URL soup = getData(url) if soup: list_items = soup.find_all("div", class_="market_listing_row market_recent_listing_row market_listing_searchresult") nameArr = pageNameRetrieve(list_items) priceArr = pageValueRetrieve(list_items) print(nameArr) print(priceArr)
额外注意事项
- Steam市场有反爬机制,频繁请求可能会被限制IP,建议添加请求间隔(比如
time.sleep(2))。 - 页面结构可能会随Steam更新变化,定期检查选择器是否有效。
内容的提问来源于stack exchange,提问作者icedid
相关产品推荐
相关产品推荐

