Scrapy选择器遇超长行失效,爬取亚马逊餐厅无法提取内容求助
解决Scrapy Selector无法处理超长script标签后内容的问题
我之前爬取大型电商页面时刚好踩过这个坑!这其实是Scrapy底层依赖的lxml库的一个已知限制:当HTML中存在单行长超过65536字符的script标签内容时,lxml的HTML解析器会触发保护机制截断后续内容,导致Selector(不管是XPath还是CSS选择器)都无法定位到超长标签之后的元素,哪怕你的表达式在浏览器里明明能匹配到。
给你三个亲测有效的修复方案,按需选择:
方案一:预处理响应,拆分超长script行
通过下载中间件在响应进入解析环节前,把超长的script内容拆分成多行,避免lxml截断。
- 新建一个下载中间件文件(比如
middlewares.py),添加以下代码:
import re from scrapy import DownloaderMiddleware class FixLongScriptMiddleware: def process_response(self, request, response, spider): # 只处理HTML类型的响应 content_type = response.headers.get('Content-Type', b'').decode('utf-8', errors='ignore') if 'text/html' in content_type: body = response.text # 匹配script标签,将其中的超长单行内容按分号换行拆分 fixed_body = re.sub( r'(<script[^>]*>)([^<]+)(</script>)', lambda match: match.group(1) + match.group(2).replace(';', ';\n') + match.group(3), body ) # 返回处理后的新响应 return response.replace(body=fixed_body.encode(response.encoding)) return response
- 在Scrapy项目的
settings.py中启用这个中间件:
DOWNLOADER_MIDDLEWARES = { 'your_project_name.middlewares.FixLongScriptMiddleware': 543, }
这个方法能从根源上解决lxml的截断问题,不影响后续正常使用Selector解析。
方案二:绕开Selector,直接用正则提取目标内容
如果只是要提取个别元素,直接操作原始响应文本是最快捷的方式,完全不受lxml解析限制:
import re def parse(self, response): # 匹配餐厅名称的h1标签 name_match = re.search( r'<h1[^>]+class="[^"]*hw-dp-restaurant-name[^"]*"[^>]*>(.*?)</h1>', response.text, re.DOTALL ) if name_match: restaurant_name = name_match.group(1).strip() # 后续处理逻辑 yield {'name': restaurant_name}
方案三:替换解析库为BeautifulSoup
BeautifulSoup的内置解析器(比如Python自带的html.parser)对超长文本的兼容性更好,不会出现截断问题:
from bs4 import BeautifulSoup def parse(self, response): # 使用html.parser解析,避免lxml的截断问题 soup = BeautifulSoup(response.text, 'html.parser') restaurant_name = soup.find('h1', class_='hw-dp-restaurant-name').get_text(strip=True) yield {'name': restaurant_name}
如果需要更精准的HTML5解析,也可以换成html5lib解析器(需要先通过pip install html5lib安装)。
内容的提问来源于stack exchange,提问作者Ayushman Koul
相关产品推荐
相关产品推荐

