You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取TripAdvisor酒店评论代码无输出,请求技术调试帮助

解决TripAdvisor酒店评论爬取无输出问题

问题原因

你代码里使用的class选择器bdYc _Q已经失效,TripAdvisor的页面元素class命名会定期更新,导致无法匹配到评论内容;另外headers里的Access-Control-*字段属于浏览器端CORS配置,作为爬虫请求不需要携带,反而可能干扰请求。

修改后的代码

import requests
from bs4 import BeautifulSoup

# 精简有效请求头
headers = {
    'accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
    'accept-encoding': 'gzip, deflate, br',
    'accept-language': 'en-US,en;q=0.5',
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
}

url = "https://www.tripadvisor.co.uk/Hotel_Review-g304141-d447407-Reviews-or10-Sigiriya_Village_Hotel-Sigiriya_Central_Province.html#REVIEWS"
req = requests.get(url, headers=headers, timeout=5)
print(req.status_code)

if req.status_code == 200:
    soup = BeautifulSoup(req.text, 'html.parser')
    # 定位单条评论的容器
    review_blocks = soup.find_all('div', attrs={'data-test-target': 'HR_CC_CARD'})
    for block in review_blocks:
        # 提取评论标题
        title = block.find('a', attrs={'data-test-target': 'review-title'})
        if title:
            print(f"评论标题: {title.get_text(strip=True)}")
        # 提取评论内容
        content = block.find('div', attrs={'data-test-target': 'review-body'})
        if content:
            print(f"评论内容: {content.get_text(strip=True)}\n")
else:
    print(f"请求失败,状态码: {req.status_code}")

补充说明

  • 如果上述代码仍无法获取内容,说明评论是通过JavaScript动态加载的,需使用selenium或playwright等工具模拟浏览器渲染后抓取。
  • 频繁爬取易触发反爬机制,建议添加请求间隔,避免IP被封禁。

内容的提问来源于stack exchange,提问作者Shiron Pereira

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 10:56:06