You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium爬取TripAdvisor失效问题求助

Fixing TripAdvisor Scraper After Page Structure Update

Looks like TripAdvisor updated their frontend DOM structure recently, which rendered all your old class-based selectors obsolete. That's why you're hitting NoSuchElementException—the elements your code was looking for no longer exist. Let's fix this by switching to more stable selectors (like data-test-target attributes, which are designed for automation/testing and less likely to change) and updating all broken parts step by step.

Here's the fully revised working code:

import csv
import time
from selenium import webdriver
import datetime
from selenium.common.exceptions import NoSuchElementException, ElementClickInterceptedException

now = datetime.datetime.now()
driver = webdriver.Chrome('chromedriver.exe')
italia = "https://www.tripadvisor.it/Attraction_Review-g657290-d2213040-Reviews-Ex_Stabilimento_Florio_delle_Tonnare_di_Favignana_e_Formica-Isola_di_Favig.html"
driver.get(italia)
place = 'Ex_Stabilimento_Florio_delle_Tonnare_di_Favignana'
lang = 'it'

def check_exists_by_xpath(xpath):
    try:
        driver.find_element_by_xpath(xpath)
    except NoSuchElementException:
        return False
    return True

for i in range(0, 2):
    try:
        # Fix 1: Updated "Read more" button selector + handle click interception
        expand_buttons = driver.find_elements_by_xpath("//button[@data-test-target='expand-review']")
        for btn in expand_buttons:
            try:
                btn.click()
                time.sleep(2)
            except ElementClickInterceptedException:
                # Bypass popups/overlays using JS click
                driver.execute_script("arguments[0].click();", btn)
                time.sleep(2)

        # Fix 2: Updated review container selector
        container = driver.find_elements_by_xpath("//div[@data-test-target='review-card']")
        num_page_items = len(container)
        
        for j in range(num_page_items):
            # Added UTF-8 encoding to handle special Italian characters
            csvFile = open(r'Italia_en.csv', 'a', encoding='utf-8')
            csvWriter = csv.writer(csvFile)
            time.sleep(1)  # Reduced unnecessary long wait

            # Fix 3: Updated rating selector (still using bubble class pattern)
            rating_a = container[j].find_element_by_xpath(
                ".//span[contains(@class, 'ui_bubble_rating')]").get_attribute("class")
            rating_b = rating_a.split("_")
            rating = rating_b[3]

            # Fix 4: Updated review content selector
            review = container[j].find_element_by_xpath(".//div[@data-test-target='review-body']").text.replace("\n", "")

            # Fix 5: Updated review title selector
            title = container[j].find_element_by_xpath(".//div[@data-test-target='review-title']").text

            # Fix 6: Updated rating date selector
            rating_date = container[j].find_element_by_xpath(".//div[@class='cRVSd']").text

            # Fix 7: Updated review link selector
            review_link = container[j].find_element_by_xpath(".//a[@data-test-target='review-title']").get_attribute('href')

            print(f"Rating: {rating}\nTitle: {title}\nReview: {review}\nDate: {rating_date}\nLink: {review_link}\n---")
            csvWriter.writerow([place, rating, title, review, rating_date, review_link, now, lang])
            csvFile.close()

        # Fix 8: Updated next page button selector + JS click to avoid interception
        next_btn = driver.find_element_by_xpath("//a[@data-test-target='pagination-next']")
        driver.execute_script("arguments[0].click();", next_btn)
        time.sleep(5)

    except NoSuchElementException as e:
        print(f"Stopped at page {i+1}: Could not find element - {e}")
        break
    except Exception as e:
        print(f"Unexpected error on page {i+1}: {e}")
        break

driver.quit()

Key Fixes Breakdown:

  • Expand Review Button: Switched to button[@data-test-target='expand-review'] and added logic to handle popups that block clicks.
  • Review Container: Replaced review-container with div[@data-test-target='review-card']—this is the current wrapper for each review.
  • Review Content/Title: Used data-test-target attributes instead of nested class chains for more reliable targeting.
  • Next Page Button: Updated to a[@data-test-target='pagination-next'] and uses JavaScript click to avoid common element interception issues.
  • CSV Encoding: Added encoding='utf-8' to prevent garbled text from Italian special characters.
  • Error Handling: Added specific exception catches for better debugging and cleaned up the browser session with driver.quit().

Pro tip: Always verify selectors manually using your browser's dev tools (F12) before running the scraper—TripAdvisor can update their structure again without warning.

内容的提问来源于stack exchange,提问作者Ignacio Aguirre

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 11:47:33