如何在Scrapy爬虫中忽略图片、指定标签及HTML属性以提升爬取速度?
Scrapy优化:忽略图片、特定标签与HTML属性的方法
一、忽略图片以提升爬取速度
Scrapy默认不会主动下载图片,除非你启用了ImagesPipeline或跟进了图片链接。要彻底避免图片请求,可通过以下方式实现:
1. 过滤图片链接(适用于CrawlSpider)
在爬虫规则中,通过LinkExtractor的deny_extensions参数排除图片格式的链接,不让Scrapy跟进这些请求:
from scrapy.spiders import CrawlSpider, Rule from scrapy.linkextractors import LinkExtractor class MyCrawlSpider(CrawlSpider): name = "example_spider" allowed_domains = ["example.com"] start_urls = ["https://example.com"] rules = ( Rule( LinkExtractor(deny_extensions=("jpg", "jpeg", "png", "gif", "bmp", "webp")), callback="parse_item", follow=True ), ) def parse_item(self, response): # 你的解析逻辑 pass
2. 用下载中间件拦截图片请求
自定义一个下载中间件,直接拦截所有图片格式的请求,避免下载:
# 在项目的middlewares.py中添加 from scrapy.exceptions import IgnoreRequest class IgnoreImagesMiddleware: def process_request(self, request, spider): image_exts = (".jpg", ".jpeg", ".png", ".gif", ".bmp", ".webp") if any(request.url.lower().endswith(ext) for ext in image_exts): raise IgnoreRequest("Skipping image download")
然后在settings.py中启用这个中间件:
DOWNLOADER_MIDDLEWARES = { "your_project_name.middlewares.IgnoreImagesMiddleware": 543, }
3. 禁用图片管道(如果已启用)
如果你的项目中配置了ImagesPipeline,直接在settings.py的ITEM_PIPELINES中注释掉对应的行即可:
# ITEM_PIPELINES = { # "scrapy.pipelines.images.ImagesPipeline": 1, # }
二、忽略特定HTML标签
如果是在解析时不想处理某些标签(比如<script>、<style>),可以用XPath直接排除;如果需要彻底移除这些标签,可借助lxml修改响应内容:
1. 解析时排除特定标签
用XPath的not()语法跳过目标标签,只提取其他内容:
def parse_item(self, response): # 提取<body>下除<script>和<style>外所有元素的文本 clean_text = response.xpath('//body//*[not(self::script or self::style)]/text()').getall() # 合并并清理文本 clean_text = " ".join([t.strip() for t in clean_text if t.strip()])
2. 彻底移除特定标签
如果需要从HTML中完全删除目标标签,可通过lxml修改响应的DOM树:
from lxml import html def parse_item(self, response): # 将响应内容转为lxml的ElementTree tree = html.fromstring(response.text) # 移除所有<script>标签 for script in tree.xpath('//script'): script.getparent().remove(script) # 移除所有<style>标签 for style in tree.xpath('//style'): style.getparent().remove(style) # 将修改后的树转回HTML字符串,生成新的响应对象 clean_html = html.tostring(tree, encoding="unicode") clean_response = response.replace(body=clean_html) # 基于清理后的响应继续解析 content = clean_response.xpath('//div[@class="main-content"]/text()').getall()
三、完全忽略HTML属性
如果需要去掉所有HTML标签的属性,同样可以通过lxml处理响应内容,或者在解析时只关注标签文本:
1. 解析时忽略属性(仅提取文本)
如果只是不需要属性,直接用XPath的text()方法提取标签内容即可,无需关注属性:
def parse_item(self, response): # 提取所有<p>标签的文本,忽略任何属性 paragraph_texts = response.xpath('//p/text()').getall()
2. 移除所有HTML属性
要彻底清除所有标签的属性,遍历DOM树并清空每个元素的属性集合:
from lxml import html def parse_item(self, response): tree = html.fromstring(response.text) # 遍历所有元素,清空属性 for element in tree.xpath('//*'): element.attrib.clear() # 生成无属性的干净HTML clean_html = html.tostring(tree, encoding="unicode") clean_response = response.replace(body=clean_html) # 后续解析基于这个干净的响应
内容的提问来源于stack exchange,提问作者Sardar
相关产品推荐
相关产品推荐

