Python Selenium WebDriver提取谷歌餐厅评论时仅读取首个评论者信息的问题排查求助
Hey there! I see exactly what's causing your script to print the first reviewer's name repeatedly—let's break this down and fix it step by step.
The Core Problem: Absolute vs. Relative XPath
Your loop to extract reviewer names uses an absolute XPath (//div[@class='TSUbDb']//a) which starts searching from the root of the entire HTML document every time. That's why it keeps grabbing the first matching element (the first reviewer's name) instead of the one inside each person block you're iterating over.
To fix this, you need to use a relative XPath by adding a dot at the start (.//), which tells Selenium to search only within the current person element.
Bonus: Fixing the "Load All Reviews" Loop
Your current loop increments total_review by 1 each time, but scrolling usually loads multiple reviews at once. This can lead to an infinite loop or not loading all reviews correctly. Instead, you should update total_review to the actual length of the newly loaded reviews each time.
Corrected Code
Here's the revised script with both fixes applied:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time driver = webdriver.Chrome() base_url = 'https://www.google.com/search?tbs=lf:1,lf_ui:9&tbm=lcl&sxsrf=AOaemvJFjYToqQmQGGnZUovsXC1CObNK1g:1633336974491&q=10+famous+restaurants+in+Dunedin&rflfq=1&num=10&sa=X&ved=2ahUKEwiTsqaxrrDzAhXe4zgGHZPODcoQjGp6BAgKEGo&biw=1280&bih=557&dpr=2#lrd=0xa82eac0dc8bdbb4b:0x4fc9070ad0f2ac70,1,,,&rlfi=hd:;si:5749134142351780976,l,CiAxMCBmYW1vdXMgcmVzdGF1cmFudHMgaW4gRHVuZWRpbiJDUjEvZ2VvL3R5cGUvZXN0YWJsaXNobWVudF9wb2kvcG9wdWxhcl93aXRoX3RvdXJpc3Rz2gENCgcI5Q8QChgFEgIIFkiDlJ7y7YCAgAhaMhAAEAEQAhgCGAQiIDEwIGZhbW91cyByZXN0YXVyYW50cyBpbiBkdW5lZGluKgQIAxACkgESaXRhbGlhbl9yZXN0YXVyYW50mgEkQ2hkRFNVaE5NRzluUzBWSlEwRm5TVU56ZW5WaFVsOUJSUkFCqgEMEAEqCCIEZm9vZCgA,y,2qOYUvKQ1C8;mv:[[-45.8349553,170.6616387],[-45.9156414,170.4803685]]' driver.get(base_url) title = driver.find_element(By.XPATH, "//div[@class='P5Bobd']").text address = driver.find_element(By.XPATH, "//div[@class='T6pBCe']").text overall_rating = driver.find_element(By.XPATH, "//div[@class='review-score-container']//span[@class='Aq14fc']").text total_reviews_text = driver.find_element(By.XPATH, "//div[@class='review-score-container']//div//div//span//span[@class='z5jxId']").text num_reviews = int(total_reviews_text.split()[0]) # Load all reviews into browser all_reviews = WebDriverWait(driver, 3).until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'div.gws-localreviews__google-review'))) total_review = len(all_reviews) while total_review < num_reviews: # Scroll to the last review to trigger loading more driver.execute_script('arguments[0].scrollIntoView(true);', all_reviews[-1]) # Wait for loading indicator to disappear WebDriverWait(driver, 5, 0.25).until_not(EC.presence_of_element_located((By.CSS_SELECTOR, 'div[class$="activityIndicator"]'))) # Refresh the list of reviews all_reviews = WebDriverWait(driver, 3).until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'div.gws-localreviews__google-review'))) total_review = len(all_reviews) # Update to actual current count time.sleep(0.5) # Optional small delay to ensure all elements load # Read and display reviewer information person_info = driver.find_elements(By.XPATH, "//div[@class='gws-localreviews__general-reviews-block']") for person in person_info: # Use relative XPath to find name within current person block name = person.find_element(By.XPATH, ".//div[@class='TSUbDb']//a").text print(name) driver.quit() # Close browser properly when done!
Key Changes Explained
- Relative XPath: Changed
//div[@class='TSUbDb']//ato.//div[@class='TSUbDb']//aso each iteration targets the name inside the specificpersonelement. - Updated Load Loop: Replaced
total_review += 1withtotal_review = len(all_reviews)to accurately track how many reviews are loaded after each scroll. - Added
driver.quit(): Ensures the browser closes cleanly after the script finishes.
This should now correctly print every reviewer's name instead of repeating the first one. Give it a try!
内容的提问来源于stack exchange,提问作者user2293224

