使用Python Selenium爬取TripAdvisor评论时遇NoSuchElementException求助
问题分析与解决方案
核心原因
你遇到的NoSuchElementException主要来自这几个问题:
- TripAdvisor前端元素结构已更新,
IRsGHoPm这个class名失效,评论内容的标签或class已更换 - 固定
time.sleep(5)的等待逻辑不可靠,网络波动时页面元素还未完全渲染就触发抓取 - 未处理网站的Cookie授权弹窗,弹窗会遮挡或干扰元素定位
- 点击"展开评论"的操作不严谨,仅点击单个按钮可能导致部分评论仍处于折叠状态
具体修复步骤
1. 验证并更新评论内容的定位器
打开目标页面按F12启动开发者工具,右键点击评论文本选择"检查",查看实际的HTML标签和class:
- 比如当前TripAdvisor的评论内容可能在
<span class="QewHAHnZ">这类标签中(以实际页面为准) - 替换代码中抓取review的xpath为新定位,示例:
review = container[j].find_element_by_xpath(".//span[@class='QewHAHnZ']").text.replace("\n", " ")
2. 用显式等待替代固定sleep,提升可靠性
固定sleep容易受网络影响,改用Selenium显式等待,确保元素加载完成后再操作:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.common.by import By # 替换原time.sleep(5),等待评论容器加载 WebDriverWait(driver, 10).until( expected_conditions.presence_of_element_located((By.XPATH, "//div[@data-reviewid]")) ) # 点击所有展开评论按钮,确保全量展开 expand_buttons = WebDriverWait(driver, 10).until( expected_conditions.presence_of_all_elements_located((By.XPATH, ".//div[contains(@data-test-target, 'expand-review')]")) ) for btn in expand_buttons: try: driver.execute_script("arguments[0].click();", btn) # 用JS点击避免元素遮挡 except: pass
3. 处理Cookie授权弹窗
多数地区访问TripAdvisor会弹出Cookie同意框,需先处理:
driver.get(url) # 等待弹窗并点击同意(xpath需根据实际弹窗调整) try: WebDriverWait(driver, 5).until( expected_conditions.element_to_be_clickable((By.XPATH, ".//button[@id='onetrust-accept-btn-handler']")) ).click() except: pass # 无弹窗则跳过
4. 完善页面切换逻辑
切换页面时用显式等待确保下一页按钮可点击,避免 stale element 问题:
# 替换原页面点击代码 try: next_btn = WebDriverWait(driver, 10).until( expected_conditions.element_to_be_clickable((By.XPATH, './/a[@class="ui_button nav next primary "]')) ) driver.execute_script("arguments[0].click();", next_btn) except NoSuchElementException: print("已到最后一页,停止爬取") break
5. 添加异常捕获,避免脚本崩溃
抓取单个评论时添加异常捕获,防止单条评论定位失败导致整个脚本终止:
for j in range(len(container)): try: rating = container[j].find_element_by_xpath(".//span[contains(@class, 'ui_bubble_rating bubble_')]").get_attribute("class").split("_")[3] title = container[j].find_element_by_xpath(".//div[contains(@data-test-target, 'review-title')]").text # 这里用更新后的review定位器 review = container[j].find_element_by_xpath(".//span[@class='QewHAHnZ']").text.replace("\n", " ") date = " ".join(dates[j].text.split(" ")[-2:]) csvWriter.writerow([date, rating, title, review]) except NoSuchElementException: print(f"第{i+1}页第{j+1}条评论抓取失败,跳过") continue
完整修复后的代码示例
import csv from selenium import webdriver from selenium.webdriver.support import expected_conditions from selenium.webdriver.support.ui import WebDriverWait from selenium.common.exceptions import NoSuchElementException from selenium.webdriver.common.by import By import time path_to_file = "E:/Desktop/Data/Reviews.csv" # 修正路径斜杠问题 pages_to_scrape = 3 url = "https://www.tripadvisor.com/Hotel_Review-g60982-d209422-Reviews-Hilton_Waikiki_Beach-Honolulu_Oahu_Hawaii.html" driver = webdriver.Chrome() driver.get(url) # 处理Cookie弹窗 try: WebDriverWait(driver, 5).until( expected_conditions.element_to_be_clickable((By.XPATH, ".//button[@id='onetrust-accept-btn-handler']")) ).click() except: pass # 用with语句自动管理文件 with open(path_to_file, 'a', encoding="utf-8", newline='') as csvFile: csvWriter = csv.writer(csvFile) # 首次运行可写入表头:csvWriter.writerow(["日期", "评分", "标题", "评论"]) for i in range(0, pages_to_scrape): # 等待评论容器加载完成 WebDriverWait(driver, 10).until( expected_conditions.presence_of_element_located((By.XPATH, "//div[@data-reviewid]")) ) # 点击所有展开评论按钮 expand_buttons = WebDriverWait(driver, 10).until( expected_conditions.presence_of_all_elements_located((By.XPATH, ".//div[contains(@data-test-target, 'expand-review')]")) ) for btn in expand_buttons: try: driver.execute_script("arguments[0].click();", btn) time.sleep(0.5) # 给展开操作留缓冲时间 except: pass container = driver.find_elements_by_xpath("//div[@data-reviewid]") dates = driver.find_elements_by_xpath(".//div[@class='_2fxQ4TOx']") for j in range(len(container)): try: rating = container[j].find_element_by_xpath(".//span[contains(@class, 'ui_bubble_rating bubble_')]").get_attribute("class").split("_")[3] title = container[j].find_element_by_xpath(".//div[contains(@data-test-target, 'review-title')]").text # 注意:此处class需根据实际页面更新 review = container[j].find_element_by_xpath(".//span[@class='QewHAHnZ']").text.replace("\n", " ") date = " ".join(dates[j].text.split(" ")[-2:]) csvWriter.writerow([date, rating, title, review]) except NoSuchElementException: print(f"第{i+1}页第{j+1}条评论抓取失败,跳过") continue # 切换到下一页 try: next_btn = WebDriverWait(driver, 10).until( expected_conditions.element_to_be_clickable((By.XPATH, './/a[@class="ui_button nav next primary "]')) ) driver.execute_script("arguments[0].click();", next_btn) except NoSuchElementException: print("已到最后一页,停止爬取") break driver.quit()
注意:代码中评论内容的classQewHAHnZ需要你根据当前页面实际值更新,TripAdvisor会不定期修改前端代码。
内容的提问来源于stack exchange,提问作者McLeodFox
相关产品推荐
相关产品推荐

