You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Scrapy爬虫中忽略图片、指定标签及HTML属性以提升爬取速度?

Scrapy优化:忽略图片、特定标签与HTML属性的方法

一、忽略图片以提升爬取速度

Scrapy默认不会主动下载图片,除非你启用了ImagesPipeline或跟进了图片链接。要彻底避免图片请求,可通过以下方式实现:

1. 过滤图片链接(适用于CrawlSpider)

在爬虫规则中,通过LinkExtractor的deny_extensions参数排除图片格式的链接,不让Scrapy跟进这些请求:

from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor

class MyCrawlSpider(CrawlSpider):
    name = "example_spider"
    allowed_domains = ["example.com"]
    start_urls = ["https://example.com"]

    rules = (
        Rule(
            LinkExtractor(deny_extensions=("jpg", "jpeg", "png", "gif", "bmp", "webp")),
            callback="parse_item",
            follow=True
        ),
    )

    def parse_item(self, response):
        # 你的解析逻辑
        pass

2. 用下载中间件拦截图片请求

自定义一个下载中间件,直接拦截所有图片格式的请求,避免下载:

# 在项目的middlewares.py中添加
from scrapy.exceptions import IgnoreRequest

class IgnoreImagesMiddleware:
    def process_request(self, request, spider):
        image_exts = (".jpg", ".jpeg", ".png", ".gif", ".bmp", ".webp")
        if any(request.url.lower().endswith(ext) for ext in image_exts):
            raise IgnoreRequest("Skipping image download")

然后在settings.py中启用这个中间件:

DOWNLOADER_MIDDLEWARES = {
    "your_project_name.middlewares.IgnoreImagesMiddleware": 543,
}

3. 禁用图片管道(如果已启用)

如果你的项目中配置了ImagesPipeline,直接在settings.py的ITEM_PIPELINES中注释掉对应的行即可:

# ITEM_PIPELINES = {
#     "scrapy.pipelines.images.ImagesPipeline": 1,
# }

二、忽略特定HTML标签

如果是在解析时不想处理某些标签(比如<script>、<style>),可以用XPath直接排除;如果需要彻底移除这些标签,可借助lxml修改响应内容:

1. 解析时排除特定标签

用XPath的not()语法跳过目标标签,只提取其他内容:

def parse_item(self, response):
    # 提取<body>下除<script>和<style>外所有元素的文本
    clean_text = response.xpath('//body//*[not(self::script or self::style)]/text()').getall()
    # 合并并清理文本
    clean_text = " ".join([t.strip() for t in clean_text if t.strip()])

2. 彻底移除特定标签

如果需要从HTML中完全删除目标标签,可通过lxml修改响应的DOM树:

from lxml import html

def parse_item(self, response):
    # 将响应内容转为lxml的ElementTree
    tree = html.fromstring(response.text)
    
    # 移除所有<script>标签
    for script in tree.xpath('//script'):
        script.getparent().remove(script)
    # 移除所有<style>标签
    for style in tree.xpath('//style'):
        style.getparent().remove(style)
    
    # 将修改后的树转回HTML字符串,生成新的响应对象
    clean_html = html.tostring(tree, encoding="unicode")
    clean_response = response.replace(body=clean_html)
    
    # 基于清理后的响应继续解析
    content = clean_response.xpath('//div[@class="main-content"]/text()').getall()

三、完全忽略HTML属性

如果需要去掉所有HTML标签的属性,同样可以通过lxml处理响应内容,或者在解析时只关注标签文本:

1. 解析时忽略属性(仅提取文本)

如果只是不需要属性,直接用XPath的text()方法提取标签内容即可,无需关注属性:

def parse_item(self, response):
    # 提取所有<p>标签的文本,忽略任何属性
    paragraph_texts = response.xpath('//p/text()').getall()

2. 移除所有HTML属性

要彻底清除所有标签的属性,遍历DOM树并清空每个元素的属性集合:

from lxml import html

def parse_item(self, response):
    tree = html.fromstring(response.text)
    
    # 遍历所有元素,清空属性
    for element in tree.xpath('//*'):
        element.attrib.clear()
    
    # 生成无属性的干净HTML
    clean_html = html.tostring(tree, encoding="unicode")
    clean_response = response.replace(body=clean_html)
    
    # 后续解析基于这个干净的响应

内容的提问来源于stack exchange,提问作者Sardar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 13:06:23