You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Steam评论Scrapy爬虫无限滚动时重复爬取第二页问题求助

解决Steam评论重复爬取的问题

问题核心在于你硬编码了userreviewscursor参数——Steam的评论翻页完全依赖这个动态变化的游标值,而非你设置的offset或页码p。固定使用同一个cursor,自然每次都会返回同一页数据。

解决步骤:

  1. 从当前页面提取下一页的cursor
    Steam会把下一页的cursor嵌入当前页面响应中,两种常见提取方式:

    • 方式一:从页面的JS变量提取
      页面里存在g_rgAppHubContentData这个JavaScript对象,其中包含下一页的cursor值,可用正则或Scrapy选择器提取。
    • 方式二:从加载更多容器的属性提取
      页面底部的div#MoreContentContainer标签带有data-cursor属性,其值就是下一页所需的cursor。
  2. 修改爬虫逻辑,动态生成下一页请求
    移除类变量page_number,改用从页面提取的cursor判断是否还有下一页,同时构造正确的请求URL。

修改后的代码示例:

import scrapy
import re
from bs4 import BeautifulSoup

class MySpider(scrapy.Spider):
    name = "MySpider"
    download_delay = 6
    start_urls = (
        'https://steamcommunity.com/app/1794680/reviews/', 
    )

    custom_settings = {
        'LOG_LEVEL': 'WARNING',
        'LOG_ENABLED': False,
        'LOG_FILE': 'logging.txt',
        'LOG_FILE_APPEND': False,
        'REQUEST_FINGERPRINTER_IMPLEMENTATION': '2.7',
        'FEEDS': {"items.json": {"format": "json", 'overwrite': True},},
    }

    def parse(self, response):
        # 处理当前页评论
        soup = BeautifulSoup(response.text, 'lxml')
        for review in soup.find_all('div', class_="apphub_UserReviewCardContent"):
            # 替换成你的评论提取逻辑
            yield {
                'content': review.find('div', class_='apphub_CardTextContent').text.strip()
            }

        # 提取下一页cursor(方式一:从JS变量提取)
        cursor_match = re.search(r'var\s+g_rgAppHubContentData\s*=\s*\{.*?"cursor"\s*:\s*"([^"]+)"', response.text)
        # 方式二:从容器属性提取,可替换上面的正则
        # cursor = response.css('#MoreContentContainer::attr(data-cursor)').get()

        # 限制爬取4页以内
        current_page = response.meta.get('page', 1)
        if cursor_match and current_page < 4:
            next_cursor = cursor_match.group(1)
            # 构造下一页URL,复用其他参数仅替换cursor
            next_url = f'https://steamcommunity.com/app/1794680/homecontent/?userreviewscursor={next_cursor}&userreviewsoffset=10&p=2&workshopitemspage=2&readytouseitemspage=2&mtxitemspage=2&itemspage=2&screenshotspage=2&videospage=2&artpage=2&allguidepage=2&webguidepage=2&integratedguidepage=2&discussionspage=2&numperpage=10&browsefilter=trendweek&browsefilter=trendweek&l=english&appHubSubSection=10&filterLanguage=default&searchText=&maxInappropriateScore=100'
            yield scrapy.Request(next_url, callback=self.parse, meta={'page': current_page + 1})

注意事项:

  • 无需手动计算offset和p参数,Steam后台主要依据userreviewscursor返回数据,其他参数保持和初始请求一致即可。
  • 使用meta传递当前页码来限制爬取数量,避免用类变量(Scrapy是异步框架,类变量会被多请求共享导致逻辑混乱)。

内容的提问来源于stack exchange,提问作者danne

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 16:16:05