You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium迭代获取数据时始终返回相同值的问题排查

问题描述

我尝试从阿尔伯塔省住宅保护公共登记系统抓取数据,代码会逐个点击搜索结果中的10条记录,打开详情弹窗后用BeautifulSoup解析文件编号,但输出始终重复显示第一条数据的编号。

代码实现

import time
import os
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

if __name__ == '__main__':
  
  print(f"Checking Browser driver...")
  os.environ['WDM_LOG'] = '0' 
  options = Options()
  options.add_argument("start-maximized")
  options.add_experimental_option("prefs", {"profile.default_content_setting_values.notifications": 1})    
  options.add_experimental_option("excludeSwitches", ["enable-automation"])
  options.add_experimental_option('excludeSwitches', ['enable-logging'])
  options.add_experimental_option('useAutomationExtension', False)
  options.add_argument('--disable-blink-features=AutomationControlled') 
  srv=Service()
  driver = webdriver.Chrome (service=srv, options=options)      
  waitWD = WebDriverWait (driver, 10)         
  
  baseLink = "https://residentialprotection.alberta.ca/public-registry/Property"
  print(f"Working for {baseLink}")  
  driver.get (baseLink)     
  waitWD.until(EC.presence_of_element_located((By.XPATH,'//input[@aria-owns="Municipality_listbox"]'))).send_keys("Calgary")
  waitWD.until(EC.presence_of_element_located((By.XPATH, '//button[@id="show-results"]'))).click() 
  time.sleep(5) 
  countElems = driver.find_elements(By.XPATH,'//tbody//tr[@role="row"]')
  print(len(countElems))
  for idx in range(len(countElems)):
    time.sleep(3)    
    elems = driver.find_elements(By.XPATH,'//tbody//tr[@role="row"]')    
    elems[idx].click()
    time.sleep(3)
    soup = BeautifulSoup (driver.page_source, 'lxml')  
    worker = soup.find("label", {"for": "FileNumber"})
    wFileNumber = worker.find_next("td").text.strip()
    print(f"{idx}: {wFileNumber}")
    closeElem = driver.find_elements(By.XPATH,'//a[@aria-label="Close"]')[-1]
    closeElem.click()
  driver.quit()

运行输出

$ python temp2.py
Checking Browser driver...
Working for https://residentialprotection.alberta.ca/public-registry/Property
10
0: 21RU3557182
1: 21RU3557182
2: 21RU3557182
3: 21RU3557182
4: 21RU3557182
5: 21RU3557182
6: 21RU3557182
7: 21RU3557182
8: 21RU3557182
9: 21RU3557182
问题原因与解决方案

核心原因

每次打开的详情弹窗并不会从DOM中被移除,只是被隐藏。当你用BeautifulSoup解析整个页面的HTML时,soup.find("label", {"for": "FileNumber"})会始终匹配第一个加载的弹窗元素,也就是第一次打开的那条记录的文件编号,导致输出重复。

解决方案

方案1:改用Selenium直接定位可见弹窗元素(推荐)

利用Selenium的等待机制,直接定位当前显示的弹窗中的文件编号元素,避免DOM中旧元素的干扰,同时比time.sleep更可靠:

替换代码中解析文件编号的部分:

# 移除BeautifulSoup相关代码,改用Selenium等待可见元素
wFileNumber = waitWD.until(EC.visibility_of_element_located(
    (By.XPATH, '//div[@role="dialog" and @aria-modal="true"]//label[@for="FileNumber"]/following-sibling::td')
)).text.strip()

方案2:如果坚持使用BeautifulSoup,定位最后一个弹窗元素

因为每次打开弹窗都会将新的弹窗元素追加到DOM末尾,所以可以通过find_all获取所有匹配的label元素,再取最后一个:

soup = BeautifulSoup(driver.page_source, 'lxml')
# 找到所有匹配的label元素,取最后一个
workers = soup.find_all("label", {"for": "FileNumber"})
wFileNumber = workers[-1].find_next("td").text.strip()

额外优化建议

  • 移除不必要的time.sleep,改用WebDriverWait等待元素状态变化,提升代码稳定性和效率;
  • 循环中无需每次重新调用find_elements,可以直接复用初始的countElems列表,但注意如果页面有动态刷新需要重新定位。

内容的提问来源于stack exchange,提问作者Rapid1898

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 09:15:12