Scrapy宽泛爬取中URL校验失效问题求助
问题排查:带查询参数的URL绕过校验存入数据库
问题描述
我正在执行宽泛爬取任务,爬取到的URL会存入数据库,已存在的URL会更新title、last_scraped等字段。为了只爬取无查询参数、无锚点的干净URL,我在Spider的parse方法里写了URL校验逻辑,但仍有带查询参数的违规URL进入数据库,且这些URL单独测试时会触发校验失败。
爬虫parse方法代码
from urllib.parse import urlparse # Retrieve links from current page links = LinkExtractor(unique=True).extract_links(response) # Follow links for link in links: if not urlparse(link.url).fragment and not urlparse(link.url).query and urlparse(link.url).scheme in {'http', 'https'}: yield scrapy.Request(link.url, callback=self.parse, errback=self.parse_errback) else: self.logger.error(f"Bad url: {link.url}")
违规URL示例
- https://licensing.edinburgh-innovations.ed.ac.uk/auth/realms/licensing.edinburgh-innovations.ed.ac.uk/protocol/openid-connect/auth?state=84bd77fcba8577b144e5ed1d8c3c6297&scope=profile+openid%20email&response_type=code&approval_prompt=auto&redirect_uri=https%3A%2F%2Flicensing.edinburgh-innovations.ed.ac.uk%2Flogin%2Foauth&client_id=e-lucid-web
- https://licensing.edinburgh-innovations.ed.ac.uk/auth/realms/licensing.edinburgh-innovations.ed.ac.uk/protocol/openid-connect/registrations?client_id=e-lucid-web&response_type=code&scope=openid+email+profile&redirect_uri=https://licensing.edinburgh-innovations.ed.ac.uk/login/oauth&kc_locale=en
- https://bookit.eca.ed.ac.uk/av/Signin.aspx?ReturnUrl=%2Fav%2Faccount
- https://www.hub.ed.ac.uk/students/login?ReturnUrl=%2fs%2fuoebs%2fevents
pipeline.py代码
class MySqlPipeline(object): def __init__(self): self.conn = mysql.connector.connect( database = os.environ.get('MYSQL_DATABASE'), password = os.environ.get('MYSQL_ROOT_PASSWORD'), user = 'root', host = 'db', port = 3306, ) # Create cursor, used to exectue SQL statements self.cur = self.conn.cursor(buffered=True) def process_item(self, item, spider): item = EdwebCrawlerItem(**item) if item['url'] =='https://www.ed.ac.uk/': return # Check if URL exists self.cur.execute(f""" SELECT id FROM {os.environ.get('MYSQL_TABLE')} WHERE url = '{item['url']}' """ ) exists = self.cur.fetchone() # UPDATE or INSERT record as appropriate if exists: self.cur.execute(f"""UPDATE {os.environ.get('MYSQL_TABLE')} SET title = %s, page_type = %s, last_modified = %s, sub_pages = %s, last_scraped = %s WHERE url = %s""", ( item['title'], item['page_type'], item['last_modified'], item['sub_pages'], datetime.datetime.today(), item['url'] )) spider.logger.error(f"Updated: {item['url']}") else: self.cur.execute(f""" INSERT INTO {os.environ.get('MYSQL_TABLE')} (url, title, page_type, last_modified, sub_pages, last_scraped) VALUES (%s, %s, %s, %s, %s, %s) """, ( item['url'], item['title'], item['page_type'], item['last_modified'], item['sub_pages'], datetime.datetime.today() )) spider.logger.error(f"inserted: {item['url']}") # Commit sql self.conn.commit() def close_spider(self, spider): self.cur.close() self.conn.close()
自定义中间件代码
class CustomRetryMiddleware(RetryMiddleware): EXCEPTIONS_TO_RETRY = ( ConnectError, ConnectionDone, ConnectionLost, TimeoutError, TCPTimedOutError, ConnectionRefusedError ) class CustomRobotsTxtMiddleware(RobotsTxtMiddleware): EXCEPTIONS_TO_LOG = ( ConnectError, ConnectionDone, ConnectionLost, TimeoutError, TCPTimedOutError, ConnectionRefusedError ) def _logerror(self, failure, request, spider): if failure.type in self.EXCEPTIONS_TO_LOG: logger.warning(f"Error downloading robots.txt: {request} {failure.value}" ) return failure
排查原因
校验逻辑仅作用于后续爬取请求,未拦截当前页面的Item入库
你写的校验逻辑只过滤了从当前页面提取的链接是否要生成新Request继续爬取,但当前页面本身的URL(也就是最终要存入数据库的Item)并没有经过校验。也就是说,当爬虫爬取到一个带查询参数的页面时,这个页面的URL会直接被封装成Item进入Pipeline,完全绕过了你在parse方法里的链接校验。Item生成逻辑未关联URL校验
推测你在parse方法中除了处理链接,还单独生成了当前页面的Item(比如提取页面标题、类型等字段),但这部分Item的URL没有经过你写的urlparse校验,直接进入了数据库操作流程。
解决方案
方案1:在生成Item前校验当前页面URL
在parse方法里,先对当前页面的URL做校验,只有符合条件的才生成Item:
def parse(self, response): # 先校验当前页面URL是否合法 parsed_current_url = urlparse(response.url) if not parsed_current_url.fragment and not parsed_current_url.query and parsed_current_url.scheme in {'http', 'https'}: # 生成当前页面的Item item = EdwebCrawlerItem() item['url'] = response.url # 填充其他字段(title、page_type等) item['title'] = response.xpath('//title/text()').get() # ...其他字段赋值 yield item else: self.logger.error(f"Current page URL invalid, skip saving: {response.url}") # 处理后续链接的爬取(原有逻辑保留) links = LinkExtractor(unique=True).extract_links(response) for link in links: parsed_link = urlparse(link.url) if not parsed_link.fragment and not parsed_link.query and parsed_link.scheme in {'http', 'https'}: yield scrapy.Request(link.url, callback=self.parse, errback=self.parse_errback) else: self.logger.error(f"Bad url: {link.url}")
方案2:在Pipeline中添加URL校验
如果不想修改Spider的parse逻辑,可以在Pipeline的process_item方法开头添加校验,直接过滤违规URL:
from urllib.parse import urlparse class MySqlPipeline(object): # ...原有__init__等方法保留 def process_item(self, item, spider): item = EdwebCrawlerItem(**item) if item['url'] =='https://www.ed.ac.uk/': return # 新增URL校验逻辑 parsed_url = urlparse(item['url']) if parsed_url.fragment or parsed_url.query or parsed_url.scheme not in {'http', 'https'}: spider.logger.error(f"Invalid URL skipped: {item['url']}") return item # 直接返回,不执行后续入库操作 # 原有的存在性检查和入库逻辑...
内容的提问来源于stack exchange,提问作者Anton Gomes
相关产品推荐
相关产品推荐

