Scrapy问题:仅爬取URL列表首个链接,无法切换至其他城市
问题分析与修复方案
核心问题:起始URL列表语法错误
你的代码中start_requests方法里的URL列表存在关键语法问题:三个URL字符串之间缺少逗号分隔。在Python中,相邻的字符串字面量会被自动拼接成一个完整字符串,导致urls列表实际上只包含一个由三个URL拼接而成的错误地址,而非三个独立的城市页面URL。这就是爬虫仅处理罗马页面的根本原因。
修复步骤
1. 修正起始URL列表格式
在每个URL末尾添加逗号,让Python识别出三个独立的列表元素:
def start_requests(self): urls = [ 'https://www.tripadvisor.it/Restaurants-g187791-Rome_Lazio.html', 'https://www.tripadvisor.it/Restaurants-g187849-Milan_Lombardy.html', 'https://www.tripadvisor.it/Restaurants-g187830-Bergamo_Province_of_Bergamo_Lombardy.html' ] for url in urls: yield scrapy.Request(url, callback=self.parse_restaurants)
2. 优化餐厅列表去重逻辑(可选)
当前代码用list(set(...))去重会打乱页面原有顺序,若需保留顺序可改为:
seen = set() all_restaurants = [] for href in response.xpath("//div[contains(@data-test,'_list_item')]//div/div/div/span/a[starts-with(@href,'/Restaurant_Review')]/@href").extract(): if href not in seen: seen.add(href) all_restaurants.append(href)
3. 完善翻页空值判断
parse_restaurant中的翻页逻辑在找不到下一页链接时会报错,可添加空值校验:
next_page_href = response.xpath('//a[@class = "nav next ui_button primary"]/@href').extract_first() if next_page_href is not None: next_page = 'https://www.tripadvisor.it' + next_page_href yield response.follow(url=next_page, callback=self.parse_restaurant)
修复后效果
修正起始URL的逗号问题后,Scrapy会自动发起三个城市页面的请求,每个页面下的餐厅会被并发爬取,最终完成罗马、米兰、贝加莫三个城市的餐厅评论全量爬取。
内容的提问来源于stack exchange,提问作者Rodolfo
相关产品推荐
相关产品推荐

