You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium WebDriver提取谷歌餐厅评论时仅读取首个评论者信息的问题排查求助

Fixing Your Google Reviews Scraper: Duplicate Name Issue

Hey there! I see exactly what's causing your script to print the first reviewer's name repeatedly—let's break this down and fix it step by step.

The Core Problem: Absolute vs. Relative XPath

Your loop to extract reviewer names uses an absolute XPath (//div[@class='TSUbDb']//a) which starts searching from the root of the entire HTML document every time. That's why it keeps grabbing the first matching element (the first reviewer's name) instead of the one inside each person block you're iterating over.

To fix this, you need to use a relative XPath by adding a dot at the start (.//), which tells Selenium to search only within the current person element.

Bonus: Fixing the "Load All Reviews" Loop

Your current loop increments total_review by 1 each time, but scrolling usually loads multiple reviews at once. This can lead to an infinite loop or not loading all reviews correctly. Instead, you should update total_review to the actual length of the newly loaded reviews each time.

Corrected Code

Here's the revised script with both fixes applied:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

driver = webdriver.Chrome()
base_url = 'https://www.google.com/search?tbs=lf:1,lf_ui:9&tbm=lcl&sxsrf=AOaemvJFjYToqQmQGGnZUovsXC1CObNK1g:1633336974491&q=10+famous+restaurants+in+Dunedin&rflfq=1&num=10&sa=X&ved=2ahUKEwiTsqaxrrDzAhXe4zgGHZPODcoQjGp6BAgKEGo&biw=1280&bih=557&dpr=2#lrd=0xa82eac0dc8bdbb4b:0x4fc9070ad0f2ac70,1,,,&rlfi=hd:;si:5749134142351780976,l,CiAxMCBmYW1vdXMgcmVzdGF1cmFudHMgaW4gRHVuZWRpbiJDUjEvZ2VvL3R5cGUvZXN0YWJsaXNobWVudF9wb2kvcG9wdWxhcl93aXRoX3RvdXJpc3Rz2gENCgcI5Q8QChgFEgIIFkiDlJ7y7YCAgAhaMhAAEAEQAhgCGAQiIDEwIGZhbW91cyByZXN0YXVyYW50cyBpbiBkdW5lZGluKgQIAxACkgESaXRhbGlhbl9yZXN0YXVyYW50mgEkQ2hkRFNVaE5NRzluUzBWSlEwRm5TVU56ZW5WaFVsOUJSUkFCqgEMEAEqCCIEZm9vZCgA,y,2qOYUvKQ1C8;mv:[[-45.8349553,170.6616387],[-45.9156414,170.4803685]]'
driver.get(base_url)

title = driver.find_element(By.XPATH, "//div[@class='P5Bobd']").text
address = driver.find_element(By.XPATH, "//div[@class='T6pBCe']").text
overall_rating = driver.find_element(By.XPATH, "//div[@class='review-score-container']//span[@class='Aq14fc']").text
total_reviews_text = driver.find_element(By.XPATH, "//div[@class='review-score-container']//div//div//span//span[@class='z5jxId']").text
num_reviews = int(total_reviews_text.split()[0])

# Load all reviews into browser
all_reviews = WebDriverWait(driver, 3).until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'div.gws-localreviews__google-review')))
total_review = len(all_reviews)

while total_review < num_reviews:
    # Scroll to the last review to trigger loading more
    driver.execute_script('arguments[0].scrollIntoView(true);', all_reviews[-1])
    # Wait for loading indicator to disappear
    WebDriverWait(driver, 5, 0.25).until_not(EC.presence_of_element_located((By.CSS_SELECTOR, 'div[class$="activityIndicator"]')))
    # Refresh the list of reviews
    all_reviews = WebDriverWait(driver, 3).until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'div.gws-localreviews__google-review')))
    total_review = len(all_reviews)  # Update to actual current count
    time.sleep(0.5)  # Optional small delay to ensure all elements load

# Read and display reviewer information
person_info = driver.find_elements(By.XPATH, "//div[@class='gws-localreviews__general-reviews-block']")
for person in person_info:
    # Use relative XPath to find name within current person block
    name = person.find_element(By.XPATH, ".//div[@class='TSUbDb']//a").text
    print(name)

driver.quit()  # Close browser properly when done!

Key Changes Explained

  • Relative XPath: Changed //div[@class='TSUbDb']//a to .//div[@class='TSUbDb']//a so each iteration targets the name inside the specific person element.
  • Updated Load Loop: Replaced total_review += 1 with total_review = len(all_reviews) to accurately track how many reviews are loaded after each scroll.
  • Added driver.quit(): Ensures the browser closes cleanly after the script finishes.

This should now correctly print every reviewer's name instead of repeating the first one. Give it a try!

内容的提问来源于stack exchange,提问作者user2293224

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 17:59:07