You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让WHILE循环后的FOR代码正常执行?Selenium问题排查

问题解决:Selenium循环滚动后无法执行后续代码

问题概述

原代码在添加WHILE循环实现页面滚动加载全部信息后,循环后的FOR循环无法执行,代码运行中断。

核心问题分析

  1. 缩进错误:获取企业元素和收集链接的代码被错误缩进在WHILE循环内部,导致循环未终止时永远不会执行后续代码。
  2. 循环终止条件失效:原代码判断"reached the end"的英文文本,但谷歌地图巴西站点使用葡萄牙语,无法定位到该元素,导致循环无限运行。
  3. 未初始化列表:links列表未提前定义,执行links.append()会直接报错。
  4. 等待时间过短:WebDriverWait仅设置1秒,极易出现元素定位超时。
  5. 元素定位问题:电话的XPATH引用了未定义的ddd变量,绝对XPATH稳定性差。

修复后的代码

import time
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC

options = Options()
options.add_argument("start-maximized")
options.add_argument('--disable-notifications')

webdriver_service = Service('C:\\webdrivers\\chromedriver.exe')
driver = webdriver.Chrome(options=options, service=webdriver_service)
# 延长等待时间至10秒,避免超时
wait = WebDriverWait(driver, 10)

url = "https://www.google.com.br/maps/search/contabilidade+balneario+camboriu/@-26.9905418,-48.6289914,15z"
driver.get(url)

# 初始化存储链接的列表
links = []

# 修正滚动循环逻辑:先滚动,再判断是否到末尾
last_height = driver.execute_script("return document.querySelector('div[role=\"main\"] div[aria-label]').scrollHeight")
while True:
    try:
        # 定位滚动容器并滚动到底部
        scroll_container = wait.until(EC.presence_of_element_located((By.XPATH, "//div[@role='main']//div[@aria-label]")))
        driver.execute_script("arguments[0].scroll(0, arguments[0].scrollHeight);", scroll_container)
        time.sleep(1)  # 等待加载新内容
        
        # 检查是否到达底部(葡萄牙语提示文本)
        wait.until(EC.visibility_of_element_located((By.XPATH, "//span[contains(text(),'chegou ao fim')]")))
        break
    except:
        # 检查滚动高度是否变化,无变化则说明已加载完毕
        new_height = driver.execute_script("return document.querySelector('div[role=\"main\"] div[aria-label]').scrollHeight")
        if new_height == last_height:
            break
        last_height = new_height

# 循环外获取所有企业元素,收集链接
classe_empresas = driver.find_elements(By.CLASS_NAME, "hfpxzc")
for empresa in classe_empresas:
    urls = empresa.get_attribute("href")
    links.append(urls)

# 遍历链接获取详情
for paginas_individuais in links:
    driver.get(paginas_individuais)
    try:
        # 定位企业名称(更稳定的方式)
        nome = wait.until(EC.visibility_of_element_located((By.CLASS_NAME, "DUwDvf lfPIob"))).text
        print(f"Nome: {nome}")
        
        # 定位地址
        endereco = wait.until(EC.visibility_of_element_located((By.XPATH, "//button[@data-item-id='address']//div[@class='Io6YTe fontBodyMedium']"))).text
        print(f"Endereco: {endereco}")
        
        # 定位电话(匹配巴西电话格式)
        tel = wait.until(EC.visibility_of_element_located((By.XPATH, "//button[@data-item-id='phone']//div[@class='Io6YTe fontBodyMedium']"))).text
        print(f"Telefone: {tel}")
    except Exception as e:
        print(f"Erro ao obter dados: {str(e)}")

driver.quit()

关键修复点说明

  • 修正缩进:将收集链接的代码移至WHILE循环外部,确保循环结束后再执行。
  • 调整终止条件:使用葡萄牙语的"chegou ao fim"作为结束提示,同时增加滚动高度对比的兜底逻辑,避免因元素定位失败导致无限循环。
  • 初始化列表:提前定义links = [],避免运行时报错。
  • 优化等待时间:将WebDriverWait超时时间设为10秒,提升元素定位成功率。
  • 稳定元素定位:替换绝对XPATH为相对定位和类名定位,移除未定义的变量,提升代码稳定性。

内容的提问来源于stack exchange,提问作者Felipe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 18:25:34