You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy选择器遇超长行失效,爬取亚马逊餐厅无法提取内容求助

解决Scrapy Selector无法处理超长script标签后内容的问题

我之前爬取大型电商页面时刚好踩过这个坑!这其实是Scrapy底层依赖的lxml库的一个已知限制:当HTML中存在单行长超过65536字符的script标签内容时,lxml的HTML解析器会触发保护机制截断后续内容,导致Selector(不管是XPath还是CSS选择器)都无法定位到超长标签之后的元素,哪怕你的表达式在浏览器里明明能匹配到。

给你三个亲测有效的修复方案,按需选择:

方案一:预处理响应,拆分超长script行

通过下载中间件在响应进入解析环节前,把超长的script内容拆分成多行,避免lxml截断。

  1. 新建一个下载中间件文件(比如middlewares.py),添加以下代码:
import re
from scrapy import DownloaderMiddleware

class FixLongScriptMiddleware:
    def process_response(self, request, response, spider):
        # 只处理HTML类型的响应
        content_type = response.headers.get('Content-Type', b'').decode('utf-8', errors='ignore')
        if 'text/html' in content_type:
            body = response.text
            # 匹配script标签,将其中的超长单行内容按分号换行拆分
            fixed_body = re.sub(
                r'(<script[^>]*>)([^<]+)(</script>)',
                lambda match: match.group(1) + match.group(2).replace(';', ';\n') + match.group(3),
                body
            )
            # 返回处理后的新响应
            return response.replace(body=fixed_body.encode(response.encoding))
        return response
  1. 在Scrapy项目的settings.py中启用这个中间件:
DOWNLOADER_MIDDLEWARES = {
    'your_project_name.middlewares.FixLongScriptMiddleware': 543,
}

这个方法能从根源上解决lxml的截断问题,不影响后续正常使用Selector解析。

方案二:绕开Selector,直接用正则提取目标内容

如果只是要提取个别元素,直接操作原始响应文本是最快捷的方式,完全不受lxml解析限制:

import re

def parse(self, response):
    # 匹配餐厅名称的h1标签
    name_match = re.search(
        r'<h1[^>]+class="[^"]*hw-dp-restaurant-name[^"]*"[^>]*>(.*?)</h1>',
        response.text,
        re.DOTALL
    )
    if name_match:
        restaurant_name = name_match.group(1).strip()
        # 后续处理逻辑
        yield {'name': restaurant_name}

方案三:替换解析库为BeautifulSoup

BeautifulSoup的内置解析器(比如Python自带的html.parser)对超长文本的兼容性更好,不会出现截断问题:

from bs4 import BeautifulSoup

def parse(self, response):
    # 使用html.parser解析,避免lxml的截断问题
    soup = BeautifulSoup(response.text, 'html.parser')
    restaurant_name = soup.find('h1', class_='hw-dp-restaurant-name').get_text(strip=True)
    yield {'name': restaurant_name}

如果需要更精准的HTML5解析,也可以换成html5lib解析器(需要先通过pip install html5lib安装)。


内容的提问来源于stack exchange,提问作者Ayushman Koul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:56:28