You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取emag.ro商品评论遇InvalidChunkLength错误,如何修改代码?

解决emag.ro评论爬取的Connection broken报错问题

报错原因分析

InvalidChunkLength错误属于HTTP分块传输编码异常,核心原因是服务器检测到请求并非真实浏览器发起,主动中断了连接,仅添加延时无法绕过这类反爬机制。

修改方案及优化代码

以下是针对性调整后的代码,解决连接中断问题同时提升爬取稳定性:

import json
import time
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

# 目标商品URL
url = "https://www.emag.ro/covor-antiderapant-negru-poliester-80-x-300-cm-c027-80x300/pd/DBY5YJMBM/?ref=sponsored_products_fill_a_b_5_3&provider=rec&recid=rec_73_c449bb3e50b63cc8f6da4a42a31af359f6cbfb3c547bc5748cb6d45501a29685_1684315709&scenario_ID=73&aid=034a897a-956c-11ed-9004-0ab644dfda7c&oid=89847310"
review_url = "https://www.emag.ro/review/get-review-listing-page?id={product_id}&page={page}"

# 完善请求头,模拟真实浏览器行为
headers = {
    'User-Agent': 'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/112.0',
    'Accept': 'application/json, text/plain, */*',
    'Accept-Language': 'en-US,en;q=0.5',
    'Referer': url,
    'DNT': '1',
    'Connection': 'keep-alive'
}

# 提取商品ID
product_id = url.split("/pd/")[1].split("/")[0]
reviews = []

# 创建带重试机制的Session,复用连接提升稳定性
session = requests.Session()
retry_strategy = Retry(
    total=3,
    backoff_factor=1,
    status_forcelist=[429, 500, 502, 503, 504]
)
adapter = HTTPAdapter(max_retries=retry_strategy)
session.mount("https://", adapter)
session.mount("http://", adapter)

page = 1
while True:
    r_url = review_url.format(product_id=product_id, page=page)
    try:
        response = session.get(r_url, headers=headers, stream=False)
        response.raise_for_status()
        data = response.json()
        
        if not data.get('reviews'):
            print(f"第{page}页无评论,停止爬取")
            break
        
        # 解析并存储评论数据
        for r in data['reviews']:
            reviews.append({
                "author": r['author']['name'],
                "date": r['date'],
                "review_text": r['content'],
                "score": r['rating']
            })
        
        print(f"成功获取第{page}页评论")
        page += 1
        # 添加随机延时,避免请求频率固定
        time.sleep(1 + time.random())
        
    except (requests.RequestException, json.JSONDecodeError) as e:
        print(f"获取第{page}页评论失败,错误:{str(e)}")
        # 连接错误时延时重试,而非直接终止
        time.sleep(3)
        continue

# 保存评论到本地文件,支持中文编码
with open('reviews.json', 'w', encoding='utf-8') as f:
    json.dump(reviews, f, indent=4, ensure_ascii=False)
print(f"共获取{len(reviews)}条评论,已保存到reviews.json")

关键修改点说明

  • 完善请求头:补充Accept、Referer等字段,让请求更贴近真实浏览器的请求特征
  • Session+重试机制:复用HTTP连接,遇到5xx、429等异常状态码自动重试,提升爬取稳定性
  • 优化异常处理:连接错误时不直接终止程序,延时后继续尝试
  • 随机延时:避免固定间隔的请求被反爬规则识别
  • 编码优化:保存文件时指定utf-8编码,避免中文乱码

内容的提问来源于stack exchange,提问作者MISHA

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 17:40:39