You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取时如何去除双引号避免CSV存储格式异常?

解决Scrapy爬取CSV时双引号解析错误问题

爬取quotes to scrape网站时,提取的文本包含双引号,存入CSV文件会触发解析错误,导致引号被转义或乱码。可以在提取阶段直接移除这些引号,具体方法如下:

核心处理逻辑

CSV默认用双引号作为字段转义符,字段内的双引号会干扰解析。我们需要在提取文本后,移除所有类型的双引号(包括常规双引号"和左右引号“”)。

方法1:字符串替换

直接用replace()方法移除引号:

text = div.css('.text::text').extract_first()
if text:
    text = text.replace('"', '').replace('“', '').replace('”', '')

方法2:正则表达式(更通用)

用正则一次性匹配所有引号类型:

import re

text = div.css('.text::text').extract_first()
if text:
    text = re.sub(r'["“”]', '', text)

修改后的完整爬虫代码

import re
import scrapy
from quotetutorial.items import QuotetutorialItem  # 替换为你的项目items路径

class QuoteSpider(scrapy.Spider):
    name = 'quotes'
    start_urls = [
        'http://quotes.toscrape.com/'
    ]
    
    def parse(self, response):
        items = QuotetutorialItem()
        allDiv = response.css('.quote')
        for div in allDiv:
            # 提取并清理引号
            text = div.css('.text::text').extract_first()
            if text:
                text = re.sub(r'["“”]', '', text)
            
            # 优化其他字段提取(用extract_first()替代extract(),避免列表存储)
            authors = div.css('.author::text').extract_first()
            aboutAuthors = div.css('.quote span a::attr(href)').extract_first()
            tags = div.css('.tags .tag::text').extract()
            
            items['storeText'] = text
            items['storeAuthors'] = authors
            items['storeAboutAuthors'] = aboutAuthors
            items['storeTags'] = tags
            
            yield items

额外优化说明

  • 用extract_first()替代extract():单个字段(文本、作者、链接)返回字符串而非列表,更适合CSV存储,避免出现[]符号。
  • 用CSS选择器::attr(href)替代XPath:写法更统一,符合Scrapy的CSS选择器习惯。
  • 增加if text判断:防止提取到空值时调用字符串方法报错。

内容的提问来源于stack exchange,提问作者Faizan Ul Haq

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 15:55:15