使用Selenium爬取TripAdvisor失效问题求助
Fixing TripAdvisor Scraper After Page Structure Update
Looks like TripAdvisor updated their frontend DOM structure recently, which rendered all your old class-based selectors obsolete. That's why you're hitting NoSuchElementException—the elements your code was looking for no longer exist. Let's fix this by switching to more stable selectors (like data-test-target attributes, which are designed for automation/testing and less likely to change) and updating all broken parts step by step.
Here's the fully revised working code:
import csv import time from selenium import webdriver import datetime from selenium.common.exceptions import NoSuchElementException, ElementClickInterceptedException now = datetime.datetime.now() driver = webdriver.Chrome('chromedriver.exe') italia = "https://www.tripadvisor.it/Attraction_Review-g657290-d2213040-Reviews-Ex_Stabilimento_Florio_delle_Tonnare_di_Favignana_e_Formica-Isola_di_Favig.html" driver.get(italia) place = 'Ex_Stabilimento_Florio_delle_Tonnare_di_Favignana' lang = 'it' def check_exists_by_xpath(xpath): try: driver.find_element_by_xpath(xpath) except NoSuchElementException: return False return True for i in range(0, 2): try: # Fix 1: Updated "Read more" button selector + handle click interception expand_buttons = driver.find_elements_by_xpath("//button[@data-test-target='expand-review']") for btn in expand_buttons: try: btn.click() time.sleep(2) except ElementClickInterceptedException: # Bypass popups/overlays using JS click driver.execute_script("arguments[0].click();", btn) time.sleep(2) # Fix 2: Updated review container selector container = driver.find_elements_by_xpath("//div[@data-test-target='review-card']") num_page_items = len(container) for j in range(num_page_items): # Added UTF-8 encoding to handle special Italian characters csvFile = open(r'Italia_en.csv', 'a', encoding='utf-8') csvWriter = csv.writer(csvFile) time.sleep(1) # Reduced unnecessary long wait # Fix 3: Updated rating selector (still using bubble class pattern) rating_a = container[j].find_element_by_xpath( ".//span[contains(@class, 'ui_bubble_rating')]").get_attribute("class") rating_b = rating_a.split("_") rating = rating_b[3] # Fix 4: Updated review content selector review = container[j].find_element_by_xpath(".//div[@data-test-target='review-body']").text.replace("\n", "") # Fix 5: Updated review title selector title = container[j].find_element_by_xpath(".//div[@data-test-target='review-title']").text # Fix 6: Updated rating date selector rating_date = container[j].find_element_by_xpath(".//div[@class='cRVSd']").text # Fix 7: Updated review link selector review_link = container[j].find_element_by_xpath(".//a[@data-test-target='review-title']").get_attribute('href') print(f"Rating: {rating}\nTitle: {title}\nReview: {review}\nDate: {rating_date}\nLink: {review_link}\n---") csvWriter.writerow([place, rating, title, review, rating_date, review_link, now, lang]) csvFile.close() # Fix 8: Updated next page button selector + JS click to avoid interception next_btn = driver.find_element_by_xpath("//a[@data-test-target='pagination-next']") driver.execute_script("arguments[0].click();", next_btn) time.sleep(5) except NoSuchElementException as e: print(f"Stopped at page {i+1}: Could not find element - {e}") break except Exception as e: print(f"Unexpected error on page {i+1}: {e}") break driver.quit()
Key Fixes Breakdown:
- Expand Review Button: Switched to
button[@data-test-target='expand-review']and added logic to handle popups that block clicks. - Review Container: Replaced
review-containerwithdiv[@data-test-target='review-card']—this is the current wrapper for each review. - Review Content/Title: Used
data-test-targetattributes instead of nested class chains for more reliable targeting. - Next Page Button: Updated to
a[@data-test-target='pagination-next']and uses JavaScript click to avoid common element interception issues. - CSV Encoding: Added
encoding='utf-8'to prevent garbled text from Italian special characters. - Error Handling: Added specific exception catches for better debugging and cleaned up the browser session with
driver.quit().
Pro tip: Always verify selectors manually using your browser's dev tools (F12) before running the scraper—TripAdvisor can update their structure again without warning.
内容的提问来源于stack exchange,提问作者Ignacio Aguirre
相关产品推荐
相关产品推荐

