You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy分页请求异常:解析评论国家码存JSON时翻页失败求助

解决Scrapy爬取速卖通评论翻页失效及国家代码提取优化问题

嘿,我看你在爬速卖通评论的国家代码时碰到翻页失效的问题了,而且代码里还有几个可以优化的地方,我帮你梳理下核心问题,再给你完整的修复方案:

一、翻页失效的核心原因

1. 全局页码变量的异步冲突

你把index定义在函数外部,而Scrapy是异步框架,多个请求会同时修改这个变量,导致页码混乱,根本无法正确请求到对应页面的内容。正确的做法是用Scrapy的response.meta来传递当前页码,让每个请求的页码独立可控。

2. 无效的CSRF Token占位符

你的formdata里的${csrfToken.parameterName}和${csrfToken.token}是模板占位符,实际请求时必须替换成页面中真实的CSRF Token值,不然速卖通服务器会直接拒绝你的请求。你可以从页面的<meta>标签或隐藏输入框中提取这个Token。

3. 缺少翻页终止条件

目前你的代码会无限发送翻页请求,直到被服务器拦截。应该判断当前页是否还有评论内容,或者是否存在下一页按钮,来停止不必要的请求。

二、国家代码提取的优化

你现在用字符串分割提取国家代码的方式太脆弱,一旦页面结构微调就会失效。建议用BeautifulSoup的原生API直接提取,更稳定:

# 假设页面结构是<div class="user-country"><b>CN</b></div>
for country_elem in soup.find_all('div', class_='user-country'):
    b_tag = country_elem.find('b')
    if b_tag:
        code = b_tag.get_text(strip=True)

三、JSON保存格式的修复

你现在每次json.dump(code, f)会把单个字符串写入文件,最终文件会变成"CN""US""JP"这种无效JSON。建议收集所有国家代码到列表,再统一写入,或者维护已有数据进行追加:

四、修改后的完整代码

def parse_fb(self, response):
    soup = BeautifulSoup(response.body, "html.parser")
    
    # 提取国家代码(优化版)
    country_codes = []
    for country_elem in soup.find_all('div', class_='user-country'):
        b_tag = country_elem.find('b')
        if b_tag:
            code = b_tag.get_text(strip=True)
            country_codes.append(code)
            print(code)
    
    # 保存到JSON(保证格式正确)
    json_file = f"{ArticlesSpider.pro_id}.json"
    if country_codes:
        # 读取已有数据(如果文件不存在则初始化空列表)
        try:
            with open(json_file, 'r') as f:
                existing_data = json.load(f)
        except FileNotFoundError:
            existing_data = []
        # 追加新数据并写入
        existing_data.extend(country_codes)
        with open(json_file, 'w') as f:
            json.dump(existing_data, f, indent=2)
    
    # 处理翻页逻辑
    # 1. 获取当前页码(首次请求默认是1)
    current_page = response.meta.get('page', 1)
    
    # 2. 提取页面中的CSRF Token(根据实际页面结构调整选择器)
    csrf_token = response.css('meta[name="csrf-token"]::attr(content)').get()
    csrf_param = response.css('meta[name="csrf-param"]::attr(content)').get()
    
    # 3. 判断是否还有下一页(比如检查当前页是否有评论)
    has_next_page = bool(soup.find_all('div', class_='user-country'))
    # 限制最大页数,避免无限请求
    if has_next_page and current_page < 10:
        request_url = 'https://feedback.aliexpress.com/display/productEvaluation.htm'
        data = {
            'ownerMemberId': '',
            'memberType': 'seller',
            'productId': str(ArticlesSpider.pro_id),
            'companyId': '',
            'evaStarFilterValue': 'all Stars',
            'evaSortValue': 'sortdefault@feedback',
            'page': str(current_page + 1),
            'currentPage': '',
            'startValidDate': '',
            'i18n': 'false',
            'withPictures': 'false',
            'withPersonalInfo': 'false',
            'withAdditionalFeedback': 'false',
            'onlyFromMyCountry': 'false',
            'version': 'evaNlpV1_2',
            'isOpened': 'true',
            'translate': 'Y',
            'jumpToTop': 'false',
        }
        # 添加真实的CSRF Token参数
        if csrf_param and csrf_token:
            data[csrf_param] = csrf_token
        
        # 传递下一页页码到下一个请求的meta中
        yield scrapy.FormRequest(
            request_url,
            formdata=data,
            callback=self.parse_fb,
            meta={'page': current_page + 1}
        )

额外提醒

速卖通有严格的反爬机制,建议在settings.py中配置合理的USER_AGENT和DOWNLOAD_DELAY(比如设置为2秒),避免IP被封禁。如果遇到动态加载的评论,可能需要结合Selenium来处理渲染后的页面。

内容的提问来源于stack exchange,提问作者許家瑄

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:46:04