You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python和Selenium从IFRAME提取目标数据?排查代码错误

问题描述

我尝试从页面https://www.bbva.com.co/personas/productos/inversion/fondos/pais.html提取指定数值,该数值位于class为iframe_base的IFRAME内(截图参考:https://i.sstatic.net/tCJCD5Hy.png)。使用Microsoft Edge WebDriver搭配Selenium编写代码后,未能提取到任何数据,求排查错误并给出正确获取数据的方法。

原代码:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.edge.service import Service
from selenium.webdriver.edge.options import Options
import time

# Configura el controlador de Edge
edge_options = Options()
# edge_options.add_argument("--headless") 
service = Service("C:/Users/PERSONAL/Downloads/msedgedriver.exe")  
driver = webdriver.Edge(service=service, options=edge_options)

# Abre la página web
driver.get("https://www.bbva.com.co/personas/productos/inversion/fondos/pais.html")  # Reemplaza con la URL real

# Espera hasta que el iframe esté presente
time.sleep(5)  # Espera 5 segundos, ajusta según sea necesario

print("Seleccionamos el IFRAME")
iframe1 = driver.find_element(By.XPATH, "//*[@id = 'content-iframe_copy']")

print("Cambiamos el foco el IFRAME")
driver.switch_to.frame(iframe1)

print("Obtener HTML del IFRAME")
html = driver.page_source
print(html)

print("Obtener el dato")
dato = driver.find_elements(By.TAG_NAME, "g")
print(dato)

driver.quit()
错误点分析
  • 硬等待不严谨:time.sleep(5)是固定时长等待,若页面加载缓慢或iframe渲染延迟,会导致切换iframe时元素未加载完成,后续操作失效。应使用Selenium的显式等待机制,确保元素就绪后再执行操作。
  • iframe定位可能偏差:你通过id='content-iframe_copy'定位iframe,但目标iframe的标识是class为iframe_base,需确认该iframe的实际定位符是否正确(比如是否存在多个iframe,或id是否为动态生成)。
  • 目标元素定位过于宽泛:使用By.TAG_NAME, "g"定位SVG元素,返回的是页面所有g标签的列表,无法直接定位到包含目标数值的具体元素。目标数值通常嵌套在SVG的文本节点或特定容器内,需要更精准的定位规则。
修正后的代码
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.edge.service import Service
from selenium.webdriver.edge.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 配置Edge驱动
edge_options = Options()
# edge_options.add_argument("--headless")  # 如需无头模式可取消注释开启
service = Service("C:/Users/PERSONAL/Downloads/msedgedriver.exe")
driver = webdriver.Edge(service=service, options=edge_options)

try:
    driver.get("https://www.bbva.com.co/personas/productos/inversion/fondos/pais.html")
    
    # 显式等待目标iframe加载完成并切换到该iframe
    # 优先通过class定位目标iframe,若需更精准可结合src等属性筛选
    WebDriverWait(driver, 10).until(
        EC.frame_to_be_available_and_switch_to_it((By.CLASS_NAME, "iframe_base"))
    )
    
    # 等待目标数值所在元素加载,根据实际DOM结构调整定位符
    # 示例:假设数值在SVG的text元素中,通过class特征定位
    target_element = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.XPATH, "//svg//text[contains(@class, 'valor')]"))
    )
    
    print(f"提取到的数值:{target_element.text}")
    
finally:
    # 切回主文档(若后续还有操作需要)
    driver.switch_to.default_content()
    driver.quit()
关键说明
  1. 显式等待:使用WebDriverWait配合expected_conditions,确保iframe和目标元素完全加载后再执行操作,避免因加载延迟导致的元素定位失败。
  2. iframe精准定位:改用By.CLASS_NAME, "iframe_base"直接定位目标iframe,更符合你描述的需求。如果页面存在多个class为iframe_base的iframe,可添加@src等属性进一步筛选,比如:By.XPATH, "//iframe[@class='iframe_base' and contains(@src, 'pais')]"。
  3. 目标元素定位调整:你需要通过浏览器开发者工具查看目标数值所在的具体元素路径,替换代码中的XPath。比如若数值在class为valor-fondo的span内,可修改为By.CLASS_NAME, "valor-fondo"。

内容的提问来源于stack exchange,提问作者AXRG

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 14:37:02