Scrapy爬虫中如何获取当前URL用于CSV文件名?(Python2.7.14/Scrapy1.5)
Scrapy获取当前URL并命名CSV文件的实现方案
当然可以!Scrapy完全支持获取当前爬取页面的URL,而且实现起来非常直观,我来给你拆解具体操作步骤和代码示例:
一、如何获取当前页面URL
在Scrapy的Spider回调函数(比如parse方法)中,每个请求都会返回一个Response对象,这个对象自带的url属性就是当前页面的完整URL。你只需要直接调用response.url就能拿到它,举个简单的例子:
def parse(self, response): # 获取当前页面的URL current_page_url = response.url print("当前爬取的页面URL:", current_page_url) # 后续的数据提取逻辑...
二、用URL命名CSV文件的注意事项
直接用原始URL作为文件名肯定不行——URL里包含/、?、&等操作系统不允许的特殊字符,所以必须先对URL进行处理,转换成合法的文件名。这里给你几种常用的处理方式:
方式1:替换特殊字符为下划线
用正则表达式把所有非字母数字、下划线、连字符、点号的字符替换成下划线,简单直接:
import re current_url = response.url # 处理成合法文件名 safe_filename = re.sub(r'[^\w\-_.]', '_', current_url) + '.csv'
方式2:生成URL的哈希值(推荐)
如果URL太长,用哈希值(比如MD5)作为文件名会更简洁,还能避免文件名过长的问题:
import hashlib current_url = response.url # 生成MD5哈希值作为文件名(Python2.7注意编码) url_hash = hashlib.md5(current_url.encode('utf-8')).hexdigest() safe_filename = "{}.csv".format(url_hash)
方式3:URL编码转义
用urllib.quote对URL进行编码,把特殊字符转成百分号编码:
import urllib current_url = response.url safe_filename = urllib.quote(current_url, safe='') + '.csv'
三、完整实现代码示例
下面给你两种常用的实现方案,根据你的需求选择:
方案1:在Spider中直接写入CSV(适合单页面生成单个CSV)
如果每个页面的数据要单独存成一个CSV文件,直接在parse方法里处理写入最直接:
import scrapy import csv import re class MySpider(scrapy.Spider): name = "my_spider" start_urls = ["http://example.com/page1", "http://example.com/page2"] def parse(self, response): # 1. 获取当前页面URL current_url = response.url # 2. 处理合法文件名 safe_filename = re.sub(r'[^\w\-_.]', '_', current_url) + '.csv' # 3. 提取页面数据(这里以提取商品信息为例) items = [] for product in response.css('div.product-item'): item = { 'title': product.css('h2.product-title::text').extract_first(), 'price': product.css('span.price::text').extract_first(), 'source_url': current_url } items.append(item) # 4. 写入CSV文件(Python2.7注意用wb模式,处理编码) if items: with open(safe_filename, 'wb') as csv_file: writer = csv.DictWriter(csv_file, fieldnames=items[0].keys()) writer.writeheader() # 处理中文编码,Python2.7需要把每个值转成utf-8 for item in items: encoded_item = {k: v.encode('utf-8') if v else v for k, v in item.items()} writer.writerow(encoded_item)
方案2:用Pipeline批量处理(适合多页面数据按URL分类存储)
如果你的爬虫会爬取多个页面,想把每个页面的数据分别写入对应URL命名的CSV,可以用Scrapy的Pipeline来实现,代码更模块化:
第一步:在Spider中把URL存入Item
import scrapy class MyItem(scrapy.Item): title = scrapy.Field() price = scrapy.Field() source_url = scrapy.Field() class MySpider(scrapy.Spider): name = "my_spider" start_urls = ["http://example.com/page1", "http://example.com/page2"] def parse(self, response): current_url = response.url for product in response.css('div.product-item'): item = MyItem() item['title'] = product.css('h2.product-title::text').extract_first() item['price'] = product.css('span.price::text').extract_first() item['source_url'] = current_url yield item
第二步:编写CSV Pipeline
import csv import re import os from scrapy.exceptions import DropItem class CSVPipeline(object): def process_item(self, item, spider): # 处理文件名 safe_filename = re.sub(r'[^\w\-_.]', '_', item['source_url']) + '.csv' # 判断文件是否存在,不存在则写入表头 file_exists = os.path.isfile(safe_filename) with open(safe_filename, 'ab') as csv_file: writer = csv.DictWriter(csv_file, fieldnames=item.fields.keys()) if not file_exists: writer.writeheader() # Python2.7处理中文编码 encoded_item = {k: v.encode('utf-8') if v else v for k, v in item.items()} writer.writerow(encoded_item) return item
第三步:在settings.py中启用Pipeline
ITEM_PIPELINES = { 'myproject.pipelines.CSVPipeline': 300, }
注意事项
- Python2.7的
csv模块对中文支持有限,写入时记得把字符串转成utf-8编码,避免乱码; - 如果同一个URL会被多次爬取,要注意处理文件覆盖问题(比如用追加模式
ab); - 哈希值方式虽然简洁,但无法通过文件名反推原始URL,如果需要保留URL关联性,优先用替换特殊字符的方式。
内容的提问来源于stack exchange,提问作者OmK
相关产品推荐
相关产品推荐

