You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy宽泛爬取中URL校验失效问题求助

问题排查:带查询参数的URL绕过校验存入数据库

问题描述

我正在执行宽泛爬取任务,爬取到的URL会存入数据库,已存在的URL会更新title、last_scraped等字段。为了只爬取无查询参数、无锚点的干净URL,我在Spider的parse方法里写了URL校验逻辑,但仍有带查询参数的违规URL进入数据库,且这些URL单独测试时会触发校验失败。

爬虫parse方法代码

from urllib.parse import urlparse

# Retrieve links from current page
links = LinkExtractor(unique=True).extract_links(response) 

# Follow links
for link in links:
    if not urlparse(link.url).fragment and not urlparse(link.url).query and urlparse(link.url).scheme in {'http', 'https'}:
        yield scrapy.Request(link.url, callback=self.parse, errback=self.parse_errback)
    else:
        self.logger.error(f"Bad url: {link.url}")

违规URL示例

  • https://licensing.edinburgh-innovations.ed.ac.uk/auth/realms/licensing.edinburgh-innovations.ed.ac.uk/protocol/openid-connect/auth?state=84bd77fcba8577b144e5ed1d8c3c6297&scope=profile+openid%20email&response_type=code&approval_prompt=auto&redirect_uri=https%3A%2F%2Flicensing.edinburgh-innovations.ed.ac.uk%2Flogin%2Foauth&client_id=e-lucid-web
  • https://licensing.edinburgh-innovations.ed.ac.uk/auth/realms/licensing.edinburgh-innovations.ed.ac.uk/protocol/openid-connect/registrations?client_id=e-lucid-web&response_type=code&scope=openid+email+profile&redirect_uri=https://licensing.edinburgh-innovations.ed.ac.uk/login/oauth&kc_locale=en
  • https://bookit.eca.ed.ac.uk/av/Signin.aspx?ReturnUrl=%2Fav%2Faccount
  • https://www.hub.ed.ac.uk/students/login?ReturnUrl=%2fs%2fuoebs%2fevents

pipeline.py代码

class MySqlPipeline(object):

    def __init__(self):
        self.conn = mysql.connector.connect(
            database = os.environ.get('MYSQL_DATABASE'),
            password = os.environ.get('MYSQL_ROOT_PASSWORD'),
            user = 'root',
            host = 'db',
            port = 3306,
        )

        # Create cursor, used to exectue SQL statements
        self.cur = self.conn.cursor(buffered=True)

    def process_item(self, item, spider):
        item = EdwebCrawlerItem(**item)
        if item['url'] =='https://www.ed.ac.uk/':
            return
        
        # Check if URL exists
        self.cur.execute(f""" SELECT id FROM {os.environ.get('MYSQL_TABLE')} WHERE url = '{item['url']}' """ )
        exists = self.cur.fetchone()

        # UPDATE or INSERT record as appropriate
        if exists:
            self.cur.execute(f"""UPDATE {os.environ.get('MYSQL_TABLE')} SET 
                title = %s,
                page_type = %s,
                last_modified = %s,
                sub_pages = %s,
                last_scraped = %s
                WHERE url = %s""", (
                item['title'],
                item['page_type'],
                item['last_modified'],
                item['sub_pages'],
                datetime.datetime.today(),
                item['url']
            ))
            spider.logger.error(f"Updated: {item['url']}")
        else:
            self.cur.execute(f""" INSERT INTO {os.environ.get('MYSQL_TABLE')} (url, title, page_type, last_modified, sub_pages, last_scraped) VALUES (%s, %s, %s, %s, %s, %s) """, (
                item['url'],
                item['title'],
                item['page_type'],
                item['last_modified'],
                item['sub_pages'],
                datetime.datetime.today()
            ))
            spider.logger.error(f"inserted: {item['url']}")

        # Commit sql
        self.conn.commit()

    def close_spider(self, spider):
        self.cur.close()
        self.conn.close()

自定义中间件代码

class CustomRetryMiddleware(RetryMiddleware):
    EXCEPTIONS_TO_RETRY = (
        ConnectError,
        ConnectionDone,
        ConnectionLost,
        TimeoutError, 
        TCPTimedOutError, 
        ConnectionRefusedError
        )

class CustomRobotsTxtMiddleware(RobotsTxtMiddleware):

    EXCEPTIONS_TO_LOG = ( 
        ConnectError,
        ConnectionDone,
        ConnectionLost,
        TimeoutError, 
        TCPTimedOutError, 
        ConnectionRefusedError
        )

    def _logerror(self, failure, request, spider):
        if failure.type in self.EXCEPTIONS_TO_LOG:
            logger.warning(f"Error downloading robots.txt: {request} {failure.value}"
            )
        return failure

排查原因

  1. 校验逻辑仅作用于后续爬取请求,未拦截当前页面的Item入库
    你写的校验逻辑只过滤了从当前页面提取的链接是否要生成新Request继续爬取,但当前页面本身的URL(也就是最终要存入数据库的Item)并没有经过校验。也就是说,当爬虫爬取到一个带查询参数的页面时,这个页面的URL会直接被封装成Item进入Pipeline,完全绕过了你在parse方法里的链接校验。

  2. Item生成逻辑未关联URL校验
    推测你在parse方法中除了处理链接,还单独生成了当前页面的Item(比如提取页面标题、类型等字段),但这部分Item的URL没有经过你写的urlparse校验,直接进入了数据库操作流程。


解决方案

方案1:在生成Item前校验当前页面URL

在parse方法里,先对当前页面的URL做校验,只有符合条件的才生成Item:

def parse(self, response):
    # 先校验当前页面URL是否合法
    parsed_current_url = urlparse(response.url)
    if not parsed_current_url.fragment and not parsed_current_url.query and parsed_current_url.scheme in {'http', 'https'}:
        # 生成当前页面的Item
        item = EdwebCrawlerItem()
        item['url'] = response.url
        # 填充其他字段(title、page_type等)
        item['title'] = response.xpath('//title/text()').get()
        # ...其他字段赋值
        yield item
    else:
        self.logger.error(f"Current page URL invalid, skip saving: {response.url}")

    # 处理后续链接的爬取(原有逻辑保留)
    links = LinkExtractor(unique=True).extract_links(response) 
    for link in links:
        parsed_link = urlparse(link.url)
        if not parsed_link.fragment and not parsed_link.query and parsed_link.scheme in {'http', 'https'}:
            yield scrapy.Request(link.url, callback=self.parse, errback=self.parse_errback)
        else:
            self.logger.error(f"Bad url: {link.url}")

方案2:在Pipeline中添加URL校验

如果不想修改Spider的parse逻辑,可以在Pipeline的process_item方法开头添加校验,直接过滤违规URL:

from urllib.parse import urlparse

class MySqlPipeline(object):
    # ...原有__init__等方法保留

    def process_item(self, item, spider):
        item = EdwebCrawlerItem(**item)
        if item['url'] =='https://www.ed.ac.uk/':
            return
        
        # 新增URL校验逻辑
        parsed_url = urlparse(item['url'])
        if parsed_url.fragment or parsed_url.query or parsed_url.scheme not in {'http', 'https'}:
            spider.logger.error(f"Invalid URL skipped: {item['url']}")
            return item  # 直接返回,不执行后续入库操作

        # 原有的存在性检查和入库逻辑...

内容的提问来源于stack exchange,提问作者Anton Gomes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 02:27:34