使用BeautifulSoup爬取Tripadvisor酒店评论详情返回空值求助
解决Tripadvisor评论详情爬取空值问题
问题根源
- 代码未导入
requests库却直接使用,且未定义请求头,可能被网站反爬拦截,导致获取的页面不完整 - URL存在空格(
or30- Pan_Pacific),请求地址无效 - 评论详情的CSS类选择器不准确,Tripadvisor的评论内容容器类名易更新,且部分评论默认折叠,需要定位正确的元素
修正后的代码
from bs4 import BeautifulSoup import requests # 修正URL空格问题 URL = "https://www.tripadvisor.jp/Hotel_Review-g294265-d302294-Reviews-or30-Pan_Pacific_Singapore-Singapore.html#REVIEWS" def get_info(link): # 添加请求头模拟浏览器,避免反爬 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(link, headers=headers) soup = BeautifulSoup(response.text, "lxml") # 用更稳定的属性定位评论容器 for items in soup.find_all("div", {"data-test-target": "review-card"}): # 提取评论者名称 name = items.find("a", {"class": "ui_header_link uyyBf"}).text.strip() # 提取评分 rate_elem = items.find("div", {"data-test-target": "review-rating"}) rate = rate_elem.find("span")["class"][1].replace("bubble_", "").strip("0") # 提取评论详情,定位正确的内容容器 details_elem = items.find("div", {"data-test-target": "review-body"}) # 处理折叠评论,获取完整内容 details = details_elem.find("span", {"class": "QewHA H4 _a"}).text.strip() if details_elem else "无评论内容" print(f"name: {name} rate: {rate}") print(f"评论详情: {details}\n") if __name__ == '__main__': get_info(URL)
关键调整说明
- 补全
requests导入和请求头,模拟正常浏览器访问,规避反爬限制 - 修正URL中的空格,确保请求地址有效
- 使用
data-test-target属性定位元素(如review-card、review-body),比CSS类更稳定,不易因网站样式更新失效 - 优化评分提取逻辑,直接通过span的class属性获取评分,无需复杂字符串处理
- 兼容评论折叠场景,定位到展开后的内容容器,确保获取完整评论
内容的提问来源于stack exchange,提问作者overNemophlia
相关产品推荐
相关产品推荐

