You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫中如何获取当前URL用于CSV文件名?(Python2.7.14/Scrapy1.5)

Scrapy获取当前URL并命名CSV文件的实现方案

当然可以!Scrapy完全支持获取当前爬取页面的URL,而且实现起来非常直观,我来给你拆解具体操作步骤和代码示例:

一、如何获取当前页面URL

在Scrapy的Spider回调函数(比如parse方法)中,每个请求都会返回一个Response对象,这个对象自带的url属性就是当前页面的完整URL。你只需要直接调用response.url就能拿到它,举个简单的例子:

def parse(self, response):
    # 获取当前页面的URL
    current_page_url = response.url
    print("当前爬取的页面URL:", current_page_url)
    # 后续的数据提取逻辑...

二、用URL命名CSV文件的注意事项

直接用原始URL作为文件名肯定不行——URL里包含/、?、&等操作系统不允许的特殊字符,所以必须先对URL进行处理,转换成合法的文件名。这里给你几种常用的处理方式:

方式1:替换特殊字符为下划线

用正则表达式把所有非字母数字、下划线、连字符、点号的字符替换成下划线,简单直接:

import re

current_url = response.url
# 处理成合法文件名
safe_filename = re.sub(r'[^\w\-_.]', '_', current_url) + '.csv'

方式2:生成URL的哈希值(推荐)

如果URL太长,用哈希值(比如MD5)作为文件名会更简洁,还能避免文件名过长的问题:

import hashlib

current_url = response.url
# 生成MD5哈希值作为文件名(Python2.7注意编码)
url_hash = hashlib.md5(current_url.encode('utf-8')).hexdigest()
safe_filename = "{}.csv".format(url_hash)

方式3:URL编码转义

用urllib.quote对URL进行编码,把特殊字符转成百分号编码:

import urllib

current_url = response.url
safe_filename = urllib.quote(current_url, safe='') + '.csv'

三、完整实现代码示例

下面给你两种常用的实现方案,根据你的需求选择:

方案1:在Spider中直接写入CSV(适合单页面生成单个CSV)

如果每个页面的数据要单独存成一个CSV文件,直接在parse方法里处理写入最直接:

import scrapy
import csv
import re

class MySpider(scrapy.Spider):
    name = "my_spider"
    start_urls = ["http://example.com/page1", "http://example.com/page2"]

    def parse(self, response):
        # 1. 获取当前页面URL
        current_url = response.url
        # 2. 处理合法文件名
        safe_filename = re.sub(r'[^\w\-_.]', '_', current_url) + '.csv'
        # 3. 提取页面数据(这里以提取商品信息为例)
        items = []
        for product in response.css('div.product-item'):
            item = {
                'title': product.css('h2.product-title::text').extract_first(),
                'price': product.css('span.price::text').extract_first(),
                'source_url': current_url
            }
            items.append(item)
        # 4. 写入CSV文件(Python2.7注意用wb模式,处理编码)
        if items:
            with open(safe_filename, 'wb') as csv_file:
                writer = csv.DictWriter(csv_file, fieldnames=items[0].keys())
                writer.writeheader()
                # 处理中文编码,Python2.7需要把每个值转成utf-8
                for item in items:
                    encoded_item = {k: v.encode('utf-8') if v else v for k, v in item.items()}
                    writer.writerow(encoded_item)

方案2:用Pipeline批量处理(适合多页面数据按URL分类存储)

如果你的爬虫会爬取多个页面,想把每个页面的数据分别写入对应URL命名的CSV,可以用Scrapy的Pipeline来实现,代码更模块化:

第一步:在Spider中把URL存入Item

import scrapy

class MyItem(scrapy.Item):
    title = scrapy.Field()
    price = scrapy.Field()
    source_url = scrapy.Field()

class MySpider(scrapy.Spider):
    name = "my_spider"
    start_urls = ["http://example.com/page1", "http://example.com/page2"]

    def parse(self, response):
        current_url = response.url
        for product in response.css('div.product-item'):
            item = MyItem()
            item['title'] = product.css('h2.product-title::text').extract_first()
            item['price'] = product.css('span.price::text').extract_first()
            item['source_url'] = current_url
            yield item

第二步:编写CSV Pipeline

import csv
import re
import os
from scrapy.exceptions import DropItem

class CSVPipeline(object):
    def process_item(self, item, spider):
        # 处理文件名
        safe_filename = re.sub(r'[^\w\-_.]', '_', item['source_url']) + '.csv'
        # 判断文件是否存在,不存在则写入表头
        file_exists = os.path.isfile(safe_filename)
        with open(safe_filename, 'ab') as csv_file:
            writer = csv.DictWriter(csv_file, fieldnames=item.fields.keys())
            if not file_exists:
                writer.writeheader()
            # Python2.7处理中文编码
            encoded_item = {k: v.encode('utf-8') if v else v for k, v in item.items()}
            writer.writerow(encoded_item)
        return item

第三步:在settings.py中启用Pipeline

ITEM_PIPELINES = {
    'myproject.pipelines.CSVPipeline': 300,
}

注意事项

  • Python2.7的csv模块对中文支持有限,写入时记得把字符串转成utf-8编码,避免乱码;
  • 如果同一个URL会被多次爬取,要注意处理文件覆盖问题(比如用追加模式ab);
  • 哈希值方式虽然简洁,但无法通过文件名反推原始URL,如果需要保留URL关联性,优先用替换特殊字符的方式。

内容的提问来源于stack exchange,提问作者OmK

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:04:59