Scrapy爬取章节页面时CSV中URL字段重复问题求助
Scrapy爬取章节URL重复问题解决
问题描述
爬取某站点的《白痴》小说章节,从第一章页面开始遍历所有章节,但输出的CSV中每行的URL字段始终是chap001.html,无法对应各章节自身的URL。
尝试过两种写法均失败:
- 用
response.css("a[href^=chap]::attr(href)").getall()[0]:所有行URL都固定为第一个chap链接 - 用列表推导式
[x for x in range(0,50)]作为列表索引:触发TypeError: list indices must be integers or slices, not list错误
原代码如下:
import scrapy class Litspider2Spider(scrapy.Spider): name = "litspider2" allowed_domains = ["www.projekt-gutenberg.org"] start_urls = ["https://www.projekt-gutenberg.org/dostojew/idiot/chap001.html"] def parse(self, response): yield {"url" : response.css("a[href^=chap]::attr(href)").getall()[x for x in range(0,50)], "text" : response.css("div.navi-gb p::text").getall()} next_chapter = response.css("a[href^=chap]::attr(href)").getall()[-1] if next_chapter is not None: next_chapter_url = "https://www.projekt-gutenberg.org/dostojew/idiot/" + next_chapter yield response.follow(next_chapter, callback=self.parse)
当前错误CSV输出:
chap001.html,"Es war gegen Ende des November, bei Tauwetter, als sich um neun Uhr morgens... chap001.html,"Der General Iwan Fjodorowitsch Jepantschin stand mitten in seinem Arbeitszimmer...
预期CSV输出:
chap001.html,"Es war gegen Ende des November, bei Tauwetter, als sich um neun Uhr morgens... chap002.html,"Der General Iwan Fjodorowitsch Jepantschin stand mitten in seinem Arbeitszimmer... ... chap050.html,....
解决方案
错误根源
- 错误地从当前页面的chap链接列表中取URL,而非直接获取当前页面自身的URL
- 列表推导式不能直接用作列表索引,属于语法错误
- 获取下一章的逻辑有误(取页面最后一个chap链接会直接跳到最后一章,而非按顺序遍历下一章),且
response.follow的URL处理冗余
修正后的代码
import scrapy class Litspider2Spider(scrapy.Spider): name = "litspider2" allowed_domains = ["www.projekt-gutenberg.org"] start_urls = ["https://www.projekt-gutenberg.org/dostojew/idiot/chap001.html"] def parse(self, response): # 获取当前页面的章节文件名(如chap001.html) current_chapter_url = response.url.split("/")[-1] # 提取当前章节文本并整理为单行字符串 chapter_text = response.css("div.navi-gb p::text").getall() full_text = " ".join([text.strip() for text in chapter_text if text.strip()]) yield { "url": current_chapter_url, "text": full_text } # 精准定位下一章链接(站点下一章按钮文本为"weiter") next_chapter = response.css("a:contains('weiter')::attr(href)").get() if next_chapter: # 用response.follow直接处理相对路径,无需手动拼接URL yield response.follow(next_chapter, callback=self.parse)
关键修改说明
- 当前章节URL获取:通过
response.url.split("/")[-1]截取当前页面URL的最后一段,直接得到对应章节的文件名,这才是当前章节的正确标识 - 文本整理:将提取的段落列表拼接为单个字符串,避免CSV输出时文本被拆分成多行导致格式混乱
- 下一章链接定位:通过
a:contains('weiter')精准找到"下一章"按钮,确保按章节顺序遍历 - 简化跳转逻辑:
response.follow支持直接使用相对路径,无需手动拼接完整URL,减少出错概率
内容的提问来源于stack exchange,提问作者Фома Ф
相关产品推荐
相关产品推荐

