You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium WebDriver:while循环无法终止问题排查

Fixing the Infinite Loop in Your Google Reviews Scraping Script

Hey there! Let's break down why your while loop won't terminate and get that review extraction working properly. Here are the key issues and actionable fixes:

Common Causes for the Infinite Loop

1. Unreliable Total Reviews Count

Your XPath for fetching the total reviews is overly complex and might be targeting the wrong element. Google's UI elements shift over time, so fragile selectors often lead to inaccurate counts that never match loaded reviews.

2. Flaky Loading Wait Condition

Using until_not(EC.presence_of_element_located(...)) for the activity indicator isn't always reliable—sometimes the indicator doesn't appear at all, or your selector is outdated, meaning the loop never gets the "done loading" signal.

3. Inefficient Scroll Trigger

Scrolling to the last review with scrollIntoView(true) might not always trigger the next batch of reviews to load, especially if the element is only partially visible.

Step-by-Step Fixes

1. Fix the Total Reviews Calculation

Replace your total reviews code with a simpler, more stable selector that targets Google's standard total reviews element:

# Get total reviews count with a reliable selector
total_reviews_element = WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.XPATH, "//div[@jsname='jVqMGc']//span[@class='z5jxId']"))
)
total_reviews_text = total_reviews_element.text
num_reviews = int(total_reviews_text.split()[0])
print(f"Total reviews to load: {num_reviews}")

2. Improve Scroll & Loading Logic

Instead of relying on the activity indicator, wait for the number of reviews to increase after scrolling. This ensures we only proceed when new reviews are actually loaded:

all_reviews = WebDriverWait(driver, 10).until(
    EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'div.gws-localreviews__google-review'))
)

while len(all_reviews) < num_reviews:
    # Scroll smoothly to the end of the last review to trigger loading
    driver.execute_script('arguments[0].scrollIntoView({behavior: "smooth", block: "end"});', all_reviews[-1])
    
    # Wait for new reviews to load (timeout after 10s if no new reviews appear)
    try:
        WebDriverWait(driver, 10).until(
            lambda d: len(d.find_elements(By.CSS_SELECTOR, 'div.gws-localreviews__google-review')) > len(all_reviews)
        )
    except TimeoutException:
        print("No more reviews to load, exiting loop.")
        break
    
    # Update the reviews list
    all_reviews = driver.find_elements(By.CSS_SELECTOR, 'div.gws-localreviews__google-review')
    print(f"Loaded {len(all_reviews)} out of {num_reviews} reviews")

3. Add Debugging Output

Printing the current review count vs the total helps you verify if the loop is making progress, or if the total count was incorrect to begin with.

Full Corrected Script

Here's the complete updated code with all fixes:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException
import time

driver = webdriver.Chrome()
base_url = 'https://www.google.com/search?tbs=lf:1,lf_ui:9&tbm=lcl&sxsrf=AOaemvJFjYToqQmQGGnZUovsXC1CObNK1g:1633336974491&q=10+famous+restaurants+in+Dunedin&rflfq=1&num=10&sa=X&ved=2ahUKEwiTsqaxrrDzAhXe4zgGHZPODcoQjGp6BAgKEGo&biw=1280&bih=557&dpr=2#lrd=0xa82eac0dc8bdbb4b:0x4fc9070ad0f2ac70,1,,,,&rlfi=hd:;si:5749134142351780976,l,CiAxMCBmYW1vdXMgcmVzdGF1cmFudHMgaW4gRHVuZWRpbiJDUjEvZ2VvL3R5cGUvZXN0YWJsaXNobWVudF9wb2kvcG9wdWxhcl93aXRoX3RvdXJpc3Rz2gENCgcI5Q8QChgFEgIIFkiDlJ7y7YCAgAhaMhAAEAEQAhgCGAQiIDEwIGZhbW91cyByZXN0YXVyYW50cyBpbiBkdW5lZGluKgQIAxACkgESaXRhbGlhbl9yZXN0YXVyYW50mgEkQ2hkRFNVaE5NRzluUzBWSlEwRm5TVU56ZW5WaFVsOUJSUkFCqgEMEAEqCCIEZm9vZCgA,y,2qOYUvKQ1C8;mv:[[-45.8349553,170.6616387],[-45.9156414,170.4803685]]'
driver.get(base_url)

# Extract basic restaurant info with explicit waits
title = WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.XPATH, "//div[@class='P5Bobd']"))
).text
address = WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.XPATH, "//div[@class='T6pBCe']"))
).text
overall_rating = WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.XPATH, "//div[@class='review-score-container']//span[@class='Aq14fc']"))
).text

# Fix total reviews count
total_reviews_element = WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.XPATH, "//div[@jsname='jVqMGc']//span[@class='z5jxId']"))
)
total_reviews_text = total_reviews_element.text
num_reviews = int(total_reviews_text.split()[0])
print(f"Total reviews: {num_reviews}")

# Load all reviews
all_reviews = WebDriverWait(driver, 10).until(
    EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'div.gws-localreviews__google-review'))
)

while len(all_reviews) < num_reviews:
    print(f"Current reviews loaded: {len(all_reviews)}")
    
    # Scroll to the end of the last review to trigger loading
    driver.execute_script('arguments[0].scrollIntoView({behavior: "smooth", block: "end"});', all_reviews[-1])
    
    # Wait for new reviews to load
    try:
        WebDriverWait(driver, 10).until(
            lambda d: len(d.find_elements(By.CSS_SELECTOR, 'div.gws-localreviews__google-review')) > len(all_reviews)
        )
    except TimeoutException:
        print("Timeout: No more reviews loaded. Exiting loop.")
        break
    
    # Update the reviews list
    all_reviews = driver.find_elements(By.CSS_SELECTOR, 'div.gws-localreviews__google-review')

print(f"Final reviews loaded: {len(all_reviews)}")
# Add your review processing logic here...

driver.quit()

Additional Notes

  • Google may limit the number of reviews loaded at once, or require clicking a "More reviews" button for large counts. If you hit this, add logic to detect and click that button when it appears.
  • Always use explicit waits (WebDriverWait) instead of time.sleep() for more reliable scraping.
  • Be mindful of Google's scraping policies—avoid sending too many requests too quickly.

内容的提问来源于stack exchange,提问作者user2293224

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 18:13:11