You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取章节页面时CSV中URL字段重复问题求助

Scrapy爬取章节URL重复问题解决

问题描述

爬取某站点的《白痴》小说章节,从第一章页面开始遍历所有章节,但输出的CSV中每行的URL字段始终是chap001.html,无法对应各章节自身的URL。

尝试过两种写法均失败:

  • 用response.css("a[href^=chap]::attr(href)").getall()[0]:所有行URL都固定为第一个chap链接
  • 用列表推导式[x for x in range(0,50)]作为列表索引:触发TypeError: list indices must be integers or slices, not list错误

原代码如下:

import scrapy

class Litspider2Spider(scrapy.Spider):
    name = "litspider2"
    allowed_domains = ["www.projekt-gutenberg.org"]
    start_urls = ["https://www.projekt-gutenberg.org/dostojew/idiot/chap001.html"]

    def parse(self, response):
        yield {"url" : response.css("a[href^=chap]::attr(href)").getall()[x for x in range(0,50)],
                "text" : response.css("div.navi-gb p::text").getall()}
            
        next_chapter = response.css("a[href^=chap]::attr(href)").getall()[-1]

        if next_chapter is not None:
            next_chapter_url = "https://www.projekt-gutenberg.org/dostojew/idiot/" + next_chapter
            yield response.follow(next_chapter, callback=self.parse)

当前错误CSV输出:

chap001.html,"Es war gegen Ende des November, bei Tauwetter, als sich um neun Uhr morgens...
chap001.html,"Der General Iwan Fjodorowitsch Jepantschin stand mitten in seinem Arbeitszimmer...

预期CSV输出:

chap001.html,"Es war gegen Ende des November, bei Tauwetter, als sich um neun Uhr morgens...
chap002.html,"Der General Iwan Fjodorowitsch Jepantschin stand mitten in seinem Arbeitszimmer...
...
chap050.html,....

解决方案

错误根源

  1. 错误地从当前页面的chap链接列表中取URL,而非直接获取当前页面自身的URL
  2. 列表推导式不能直接用作列表索引,属于语法错误
  3. 获取下一章的逻辑有误(取页面最后一个chap链接会直接跳到最后一章,而非按顺序遍历下一章),且response.follow的URL处理冗余

修正后的代码

import scrapy

class Litspider2Spider(scrapy.Spider):
    name = "litspider2"
    allowed_domains = ["www.projekt-gutenberg.org"]
    start_urls = ["https://www.projekt-gutenberg.org/dostojew/idiot/chap001.html"]

    def parse(self, response):
        # 获取当前页面的章节文件名(如chap001.html)
        current_chapter_url = response.url.split("/")[-1]
        # 提取当前章节文本并整理为单行字符串
        chapter_text = response.css("div.navi-gb p::text").getall()
        full_text = " ".join([text.strip() for text in chapter_text if text.strip()])
        
        yield {
            "url": current_chapter_url,
            "text": full_text
        }
        
        # 精准定位下一章链接(站点下一章按钮文本为"weiter")
        next_chapter = response.css("a:contains('weiter')::attr(href)").get()
        if next_chapter:
            # 用response.follow直接处理相对路径,无需手动拼接URL
            yield response.follow(next_chapter, callback=self.parse)

关键修改说明

  • 当前章节URL获取:通过response.url.split("/")[-1]截取当前页面URL的最后一段,直接得到对应章节的文件名,这才是当前章节的正确标识
  • 文本整理:将提取的段落列表拼接为单个字符串,避免CSV输出时文本被拆分成多行导致格式混乱
  • 下一章链接定位:通过a:contains('weiter')精准找到"下一章"按钮,确保按章节顺序遍历
  • 简化跳转逻辑:response.follow支持直接使用相对路径,无需手动拼接完整URL,减少出错概率

内容的提问来源于stack exchange,提问作者Фома Ф

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 23:15:34