You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取TripAdvisor API页面无结果问题求助

解决Scrapy请求TripAdvisor API返回空结果的问题

我遇到了一个棘手的问题:手里有个TripAdvisor的TypeAhead API链接,在浏览器里换不同的query参数访问都能拿到正常的JSON响应,但用Scrapy或者Scrapy Shell请求时,要么触发Forbidden by robots.txt,要么返回空的结果({"normalized":{"query":""},"query":{},"results":[],"partial_content":false})。

我的Scrapy爬虫代码片段

link = "https://www.tripadvisor.com/TypeAheadJson?action=API&types=geo%2Cnbrhd%2Chotel%2Ctheme_park&legacy_format=true&urlList=true&strictParent=true&query={}%20dubai%20hotel&max=6&name_depth=3&interleaved=true&scoreThreshold=0.5&strictAnd=false&typeahead1_5=true&disableMaxGroupSize=true&geoBoostFix=true&neighborhood_geos=true&details=true&link_type=hotel%2Cvr%2Ceat%2Cattr&rescue=true&uiOrigin=trip_search_Hotels&source=trip_search_Hotels&startTime=1516800919604&searchSessionId=BA939B3D93510DABB510328CBF3353131516800881576ssid&nearPages=true"

def start_requests(self):
    files = [f for f in listdir('results/') if isfile(join('results/', f))]
    for file in files:
        with open('results/' + file, 'r', encoding="utf8") as tour_info:
            tour = json.load(tour_info)
            for hotel in tour["hotels"]:
                yield scrapy.Request(self.link.format(hotel))

name = 'tripadvisor'
allowed_domains = ['tripadvisor.com']

def parse(self, response):
    print(response.body)

问题细节

  • 一开始每个请求都触发Forbidden by robots.txt,把Scrapy的ROBOTSTXT_OBEY设为False后,这个错误消失了,但返回的结果是空的,完全没有预期的酒店数据。
  • 浏览器里访问时,能拿到类似这样的有效JSON:

[
{
"urls":[
{
"url_type":"hotel",
"name":"Sadaf Hotel, Dubai, United Arab Emirates",
"type":"HOTEL",
"url":"/Hotel_Review-g295424-d633008-Reviews-Sadaf_Hotel-Dubai_Emirate_of_Dubai.html"
}
],
...
}
]


解决方案建议

从你的情况来看,TripAdvisor的API在做反爬校验,主要针对非浏览器请求,下面是几个可以尝试的修复方向:

1. 模拟浏览器的请求头

TripAdvisor会检查请求的User-Agent,默认Scrapy的User-Agent是Scrapy/VERSION (+https://scrapy.org),很容易被识别。你需要把User-Agent改成浏览器的标识,比如Chrome的:

yield scrapy.Request(
    self.link.format(hotel),
    headers={
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    }
)

另外,还可以加上Accept、Accept-Language这些浏览器常用的请求头,进一步模拟真实请求。

2. 动态生成searchSessionId和startTime参数

你链接里的searchSessionId和startTime是固定的,这两个参数是TripAdvisor用来标识会话的动态参数,浏览器每次请求都会生成新的。固定值很容易被识别为爬虫,建议:

  • startTime可以用当前时间戳(毫秒级)生成,比如int(time.time() * 1000)
  • searchSessionId可以参考浏览器生成的格式,用随机十六进制字符串拼接时间戳模拟,或者直接从浏览器请求里抓取最新的会话ID复用。

3. 处理Cookie

浏览器访问时会有TripAdvisor设置的Cookie,Scrapy默认不会携带这些Cookie,你可以从浏览器DevTools里复制Cookie,加到请求头里:

headers={
    'User-Agent': '...',
    'Cookie': '你的浏览器Cookie内容'
}

或者使用Scrapy的CookieMiddleware自动处理会话Cookie,先请求一次TripAdvisor主页,获取Cookie后再请求API。

4. 检查URL编码问题

你链接里的&是HTML实体编码,在Python里应该直接用&,虽然Scrapy可能会自动处理,但最好替换一下,避免参数解析错误:把链接里的&全部替换成&,确保每个参数都能被正确识别。

5. 尝试使用无头浏览器渲染页面

如果上面的方法都不行,说明TripAdvisor的反爬机制比较严格,可以用无头浏览器模拟真实浏览器请求,比如scrapy-playwright,它可以完全模拟浏览器行为,包括JS渲染和会话管理。


内容的提问来源于stack exchange,提问作者Amirition

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:21:08