You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy问题:仅爬取URL列表首个链接,无法切换至其他城市

问题分析与修复方案

核心问题:起始URL列表语法错误

你的代码中start_requests方法里的URL列表存在关键语法问题:三个URL字符串之间缺少逗号分隔。在Python中,相邻的字符串字面量会被自动拼接成一个完整字符串,导致urls列表实际上只包含一个由三个URL拼接而成的错误地址,而非三个独立的城市页面URL。这就是爬虫仅处理罗马页面的根本原因。

修复步骤

1. 修正起始URL列表格式

在每个URL末尾添加逗号,让Python识别出三个独立的列表元素:

def start_requests(self):
    urls = [
        'https://www.tripadvisor.it/Restaurants-g187791-Rome_Lazio.html',
        'https://www.tripadvisor.it/Restaurants-g187849-Milan_Lombardy.html',
        'https://www.tripadvisor.it/Restaurants-g187830-Bergamo_Province_of_Bergamo_Lombardy.html'
    ]
    for url in urls:
        yield scrapy.Request(url, callback=self.parse_restaurants)

2. 优化餐厅列表去重逻辑(可选)

当前代码用list(set(...))去重会打乱页面原有顺序,若需保留顺序可改为:

seen = set()
all_restaurants = []
for href in response.xpath("//div[contains(@data-test,'_list_item')]//div/div/div/span/a[starts-with(@href,'/Restaurant_Review')]/@href").extract():
    if href not in seen:
        seen.add(href)
        all_restaurants.append(href)

3. 完善翻页空值判断

parse_restaurant中的翻页逻辑在找不到下一页链接时会报错,可添加空值校验:

next_page_href = response.xpath('//a[@class = "nav next ui_button primary"]/@href').extract_first()
if next_page_href is not None:
    next_page = 'https://www.tripadvisor.it' + next_page_href
    yield response.follow(url=next_page, callback=self.parse_restaurant)

修复后效果

修正起始URL的逗号问题后,Scrapy会自动发起三个城市页面的请求,每个页面下的餐厅会被并发爬取,最终完成罗马、米兰、贝加莫三个城市的餐厅评论全量爬取。

内容的提问来源于stack exchange,提问作者Rodolfo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 09:40:26