Scrapy爬取TripAdvisor巴厘岛酒店评论无输出问题求助
问题排查与修复方案
你的Scrapy爬虫输出为空,主要是以下几个关键问题导致的,逐一修复即可:
1. 域名权限配置错误
allowed_domains设置为['tripadvisor.com'],但目标站点是秘鲁站tripadvisor.com.pe,域名不匹配会导致爬虫被限制跟进任何链接。
修复:将allowed_domains改为:
allowed_domains = ['tripadvisor.com.pe']
2. 评论容器选择器过于宽泛
opiniones = sel.xpath('//div[@id="content"]/div/div')匹配了页面中大量无关的div元素,根本没有定位到用户评论节点。
修复:使用个人主页中精准的评论容器选择器:
opiniones = sel.xpath('//div[@data-test-target="review-card"]')
3. XPath语法错误与路径问题
- 标题XPath缺少
@符号(class前需要加@),且使用绝对路径会导致每次取到页面第一个标题,而非当前评论的标题; - 评分元素的XPath定位错误,个人主页的评分元素结构与酒店评论页不同。
修复: - 标题XPath改为相对路径并修正语法:
.//div[@class="AzIrY b _a VrCoN"]/text() - 评分XPath改为个人主页对应的元素:
.//div[contains(@class, "ui_bubble_rating")]/@class
4. 用户主页链接的提取范围未限制
原代码注释掉了restrict_xpaths,可能导致抓取到非评论作者的/profile/链接,浪费爬取资源且无法获取有效评论。
修复:启用并修正restrict_xpaths:
restrict_xpaths=['//a[@class="ui_header_link uyyBf"]']
修正后的完整代码
from scrapy.item import Field from scrapy.item import Item from scrapy.spiders import CrawlSpider, Rule from scrapy.selector import Selector from scrapy.loader.processors import MapCompose from scrapy.linkextractors import LinkExtractor from scrapy.loader import ItemLoader class Opinion(Item): titulo = Field() calificacion = Field() contenido = Field() autor = Field() class TripAdvisor(CrawlSpider): name = "OpinionesTripAdvisor" custom_settings = { 'USER_AGENT': "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/80.0.3987.149 Safari/537.36", 'CLOSESPIDER_PAGECOUNT':100 } download_delay = 1 allowed_domains = ['tripadvisor.com.pe'] start_urls = ["https://www.tripadvisor.com.pe/Hotels-g294226-Bali-Hotels.html"] rules = ( # 酒店列表分页 Rule( LinkExtractor( allow=r'-oa\d+-' ), follow=True ), # 酒店详情页 Rule( LinkExtractor( allow=r'/Hotel_Review-', restrict_xpaths=['//div[@id="taplc_hsx_hotel_list_lite_dusty_hotels_combined_sponsored_ad_density_control_0"]//a[data-clicksource="HotelName"]'] ), follow=True ), # 评论分页 Rule( LinkExtractor( allow=r'/-or\d+-/' ), follow=True ), # 用户个人主页,提取评论 Rule( LinkExtractor( allow=r'/profile/', restrict_xpaths=['//a[@class="ui_header_link uyyBf"]'] ), follow=True, callback='parse_opinion' ) ) # 提取评分 def obtenerCalif(self, texto): calificacion = texto.split("_")[-1] return calificacion # 解析用户评论 def parse_opinion(self, response): sel = Selector(response) opiniones = sel.xpath('//div[@data-test-target="review-card"]') autor = sel.xpath('//h1/text()').get() for opinion in opiniones: item = ItemLoader(Opinion(), opinion) item.add_value('autor', autor) item.add_xpath('titulo', './/div[@class="AzIrY b _a VrCoN"]/text()') item.add_xpath('contenido', './/q/text()') item.add_xpath('calificacion', './/div[contains(@class, "ui_bubble_rating")]/@class', MapCompose(self.obtenerCalif)) yield item.load_item()
额外注意事项
- 建议开启Scrapy的日志功能(
LOG_LEVEL='DEBUG'),可以看到链接跟进、元素匹配的详细过程,方便排查问题; - TripAdvisor有反爬机制,若爬取量较大,建议增加
download_delay,或使用代理IP池; - 页面元素类名可能会随网站更新变化,若后续再次出现匹配问题,需重新检查页面结构。
内容的提问来源于stack exchange,提问作者drone2700
相关产品推荐
相关产品推荐

