如何让WHILE循环后的FOR代码正常执行?Selenium问题排查
问题解决:Selenium循环滚动后无法执行后续代码
问题概述
原代码在添加WHILE循环实现页面滚动加载全部信息后,循环后的FOR循环无法执行,代码运行中断。
核心问题分析
- 缩进错误:获取企业元素和收集链接的代码被错误缩进在WHILE循环内部,导致循环未终止时永远不会执行后续代码。
- 循环终止条件失效:原代码判断"reached the end"的英文文本,但谷歌地图巴西站点使用葡萄牙语,无法定位到该元素,导致循环无限运行。
- 未初始化列表:
links列表未提前定义,执行links.append()会直接报错。 - 等待时间过短:
WebDriverWait仅设置1秒,极易出现元素定位超时。 - 元素定位问题:电话的XPATH引用了未定义的
ddd变量,绝对XPATH稳定性差。
修复后的代码
import time from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC options = Options() options.add_argument("start-maximized") options.add_argument('--disable-notifications') webdriver_service = Service('C:\\webdrivers\\chromedriver.exe') driver = webdriver.Chrome(options=options, service=webdriver_service) # 延长等待时间至10秒,避免超时 wait = WebDriverWait(driver, 10) url = "https://www.google.com.br/maps/search/contabilidade+balneario+camboriu/@-26.9905418,-48.6289914,15z" driver.get(url) # 初始化存储链接的列表 links = [] # 修正滚动循环逻辑:先滚动,再判断是否到末尾 last_height = driver.execute_script("return document.querySelector('div[role=\"main\"] div[aria-label]').scrollHeight") while True: try: # 定位滚动容器并滚动到底部 scroll_container = wait.until(EC.presence_of_element_located((By.XPATH, "//div[@role='main']//div[@aria-label]"))) driver.execute_script("arguments[0].scroll(0, arguments[0].scrollHeight);", scroll_container) time.sleep(1) # 等待加载新内容 # 检查是否到达底部(葡萄牙语提示文本) wait.until(EC.visibility_of_element_located((By.XPATH, "//span[contains(text(),'chegou ao fim')]"))) break except: # 检查滚动高度是否变化,无变化则说明已加载完毕 new_height = driver.execute_script("return document.querySelector('div[role=\"main\"] div[aria-label]').scrollHeight") if new_height == last_height: break last_height = new_height # 循环外获取所有企业元素,收集链接 classe_empresas = driver.find_elements(By.CLASS_NAME, "hfpxzc") for empresa in classe_empresas: urls = empresa.get_attribute("href") links.append(urls) # 遍历链接获取详情 for paginas_individuais in links: driver.get(paginas_individuais) try: # 定位企业名称(更稳定的方式) nome = wait.until(EC.visibility_of_element_located((By.CLASS_NAME, "DUwDvf lfPIob"))).text print(f"Nome: {nome}") # 定位地址 endereco = wait.until(EC.visibility_of_element_located((By.XPATH, "//button[@data-item-id='address']//div[@class='Io6YTe fontBodyMedium']"))).text print(f"Endereco: {endereco}") # 定位电话(匹配巴西电话格式) tel = wait.until(EC.visibility_of_element_located((By.XPATH, "//button[@data-item-id='phone']//div[@class='Io6YTe fontBodyMedium']"))).text print(f"Telefone: {tel}") except Exception as e: print(f"Erro ao obter dados: {str(e)}") driver.quit()
关键修复点说明
- 修正缩进:将收集链接的代码移至WHILE循环外部,确保循环结束后再执行。
- 调整终止条件:使用葡萄牙语的"chegou ao fim"作为结束提示,同时增加滚动高度对比的兜底逻辑,避免因元素定位失败导致无限循环。
- 初始化列表:提前定义
links = [],避免运行时报错。 - 优化等待时间:将
WebDriverWait超时时间设为10秒,提升元素定位成功率。 - 稳定元素定位:替换绝对XPATH为相对定位和类名定位,移除未定义的变量,提升代码稳定性。
内容的提问来源于stack exchange,提问作者Felipe
相关产品推荐
相关产品推荐

