You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Selenium爬取TripAdvisor评论时遇NoSuchElementException求助

问题分析与解决方案

核心原因

你遇到的NoSuchElementException主要来自这几个问题:

  • TripAdvisor前端元素结构已更新,IRsGHoPm这个class名失效,评论内容的标签或class已更换
  • 固定time.sleep(5)的等待逻辑不可靠,网络波动时页面元素还未完全渲染就触发抓取
  • 未处理网站的Cookie授权弹窗,弹窗会遮挡或干扰元素定位
  • 点击"展开评论"的操作不严谨,仅点击单个按钮可能导致部分评论仍处于折叠状态

具体修复步骤

1. 验证并更新评论内容的定位器

打开目标页面按F12启动开发者工具,右键点击评论文本选择"检查",查看实际的HTML标签和class:

  • 比如当前TripAdvisor的评论内容可能在<span class="QewHAHnZ">这类标签中(以实际页面为准)
  • 替换代码中抓取review的xpath为新定位,示例:
    review = container[j].find_element_by_xpath(".//span[@class='QewHAHnZ']").text.replace("\n", "  ")
    

2. 用显式等待替代固定sleep,提升可靠性

固定sleep容易受网络影响,改用Selenium显式等待,确保元素加载完成后再操作:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By

# 替换原time.sleep(5),等待评论容器加载
WebDriverWait(driver, 10).until(
    expected_conditions.presence_of_element_located((By.XPATH, "//div[@data-reviewid]"))
)

# 点击所有展开评论按钮,确保全量展开
expand_buttons = WebDriverWait(driver, 10).until(
    expected_conditions.presence_of_all_elements_located((By.XPATH, ".//div[contains(@data-test-target, 'expand-review')]"))
)
for btn in expand_buttons:
    try:
        driver.execute_script("arguments[0].click();", btn)  # 用JS点击避免元素遮挡
    except:
        pass

3. 处理Cookie授权弹窗

多数地区访问TripAdvisor会弹出Cookie同意框,需先处理:

driver.get(url)
# 等待弹窗并点击同意(xpath需根据实际弹窗调整)
try:
    WebDriverWait(driver, 5).until(
        expected_conditions.element_to_be_clickable((By.XPATH, ".//button[@id='onetrust-accept-btn-handler']"))
    ).click()
except:
    pass  # 无弹窗则跳过

4. 完善页面切换逻辑

切换页面时用显式等待确保下一页按钮可点击,避免 stale element 问题:

# 替换原页面点击代码
try:
    next_btn = WebDriverWait(driver, 10).until(
        expected_conditions.element_to_be_clickable((By.XPATH, './/a[@class="ui_button nav next primary "]'))
    )
    driver.execute_script("arguments[0].click();", next_btn)
except NoSuchElementException:
    print("已到最后一页,停止爬取")
    break

5. 添加异常捕获,避免脚本崩溃

抓取单个评论时添加异常捕获,防止单条评论定位失败导致整个脚本终止:

for j in range(len(container)):
    try:
        rating = container[j].find_element_by_xpath(".//span[contains(@class, 'ui_bubble_rating bubble_')]").get_attribute("class").split("_")[3]
        title = container[j].find_element_by_xpath(".//div[contains(@data-test-target, 'review-title')]").text
        # 这里用更新后的review定位器
        review = container[j].find_element_by_xpath(".//span[@class='QewHAHnZ']").text.replace("\n", "  ")
        date = " ".join(dates[j].text.split(" ")[-2:])
        csvWriter.writerow([date, rating, title, review])
    except NoSuchElementException:
        print(f"第{i+1}页第{j+1}条评论抓取失败,跳过")
        continue

完整修复后的代码示例

import csv
from selenium import webdriver
from selenium.webdriver.support import expected_conditions
from selenium.webdriver.support.ui import WebDriverWait
from selenium.common.exceptions import NoSuchElementException
from selenium.webdriver.common.by import By
import time

path_to_file = "E:/Desktop/Data/Reviews.csv"  # 修正路径斜杠问题
pages_to_scrape = 3

url = "https://www.tripadvisor.com/Hotel_Review-g60982-d209422-Reviews-Hilton_Waikiki_Beach-Honolulu_Oahu_Hawaii.html"

driver = webdriver.Chrome()
driver.get(url)

# 处理Cookie弹窗
try:
    WebDriverWait(driver, 5).until(
        expected_conditions.element_to_be_clickable((By.XPATH, ".//button[@id='onetrust-accept-btn-handler']"))
    ).click()
except:
    pass

# 用with语句自动管理文件
with open(path_to_file, 'a', encoding="utf-8", newline='') as csvFile:
    csvWriter = csv.writer(csvFile)
    # 首次运行可写入表头:csvWriter.writerow(["日期", "评分", "标题", "评论"])

    for i in range(0, pages_to_scrape):
        # 等待评论容器加载完成
        WebDriverWait(driver, 10).until(
            expected_conditions.presence_of_element_located((By.XPATH, "//div[@data-reviewid]"))
        )

        # 点击所有展开评论按钮
        expand_buttons = WebDriverWait(driver, 10).until(
            expected_conditions.presence_of_all_elements_located((By.XPATH, ".//div[contains(@data-test-target, 'expand-review')]"))
        )
        for btn in expand_buttons:
            try:
                driver.execute_script("arguments[0].click();", btn)
                time.sleep(0.5)  # 给展开操作留缓冲时间
            except:
                pass

        container = driver.find_elements_by_xpath("//div[@data-reviewid]")
        dates = driver.find_elements_by_xpath(".//div[@class='_2fxQ4TOx']")
        
        for j in range(len(container)):
            try:
                rating = container[j].find_element_by_xpath(".//span[contains(@class, 'ui_bubble_rating bubble_')]").get_attribute("class").split("_")[3]
                title = container[j].find_element_by_xpath(".//div[contains(@data-test-target, 'review-title')]").text
                # 注意:此处class需根据实际页面更新
                review = container[j].find_element_by_xpath(".//span[@class='QewHAHnZ']").text.replace("\n", "  ")
                date = " ".join(dates[j].text.split(" ")[-2:])
                csvWriter.writerow([date, rating, title, review])
            except NoSuchElementException:
                print(f"第{i+1}页第{j+1}条评论抓取失败,跳过")
                continue
        
        # 切换到下一页
        try:
            next_btn = WebDriverWait(driver, 10).until(
                expected_conditions.element_to_be_clickable((By.XPATH, './/a[@class="ui_button nav next primary "]'))
            )
            driver.execute_script("arguments[0].click();", next_btn)
        except NoSuchElementException:
            print("已到最后一页,停止爬取")
            break

driver.quit()

注意:代码中评论内容的classQewHAHnZ需要你根据当前页面实际值更新,TripAdvisor会不定期修改前端代码。

内容的提问来源于stack exchange,提问作者McLeodFox

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 01:15:39