使用Selenium定位SVG元素实现股票财务数据下载自动化求助
问题
需要自动化从指定网站下载股票财务数据,以已订阅可访问的Turkiye Sigorta为例,目标URL:https://fintables.com/sirketler/TURSG/finansal-tablolar/gelir-tablosu?period=&type=¤cy=usd
下载Excel格式数据的按钮文本“Excel'e Aktar”位于SVG元素内,以下是无法正常运行的Selenium代码:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import NoSuchElementException, TimeoutException url = 'https://fintables.com/sirketler/TURSG/finansal-tablolar/bilanco?period=&type=¤cy=usd' driver = webdriver.Chrome() driver.maximize_window() driver.get(url) try: svg_element = WebDriverWait(driver,10).until(EC.presence_of_element_located((By.XPATH,"//*[name()='svg']"))) # Find the specific SVG element containing the text 'Excel'e Aktar' (with capital 'A') excel_button_svg = svg_element.find_element(By.XPATH, f".//*[contains(text(), \"Excel'e Aktar\")]") # Click the 'Excel'e Aktar' button within the SVG excel_button_svg.click() except NoSuchElementException: print("SVG or button not found") except TimeoutException: print("Timed out waiting for SVG element to be clickable") except Exception as e: print(f"An unexpected error occurred: {str(e)}")
修复后需将流程扩展至550+只股票的批量下载任务。
解决方案
原代码核心问题
- SVG定位范围过宽:直接匹配所有SVG元素,大概率定位到无关的SVG节点,而非下载按钮对应的目标元素。
- 文本匹配逻辑错误:SVG内的文本通常嵌套在
<text>或<tspan>标签中,原代码的XPATH无法精准匹配;且先找单个SVG再内部查找文本的逻辑,会遗漏按钮容器(一般是<button>或<a>标签包裹SVG)。 - 等待条件不适用:
presence_of_element_located仅确保元素存在,未考虑元素是否可点击,容易导致点击失败。
修复后的代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import NoSuchElementException, TimeoutException url = 'https://fintables.com/sirketler/TURSG/finansal-tablolar/bilanco?period=&type=¤cy=usd' driver = webdriver.Chrome() driver.maximize_window() driver.get(url) try: # 直接定位包裹目标SVG的可点击按钮,精准匹配SVG内的文本标签 excel_button = WebDriverWait(driver, 15).until( EC.element_to_be_clickable( (By.XPATH, "//button[.//*[name()='svg']//*[name()='text' and contains(text(), \"Excel'e Aktar\")]]") ) ) excel_button.click() print("Excel下载按钮点击成功") except NoSuchElementException: print("未找到下载按钮或对应的SVG元素") except TimeoutException: print("等待下载按钮可点击超时") except Exception as e: print(f"发生未预期错误: {str(e)}") finally: # 批量任务时可保留浏览器实例,避免重复启动 # driver.quit() pass
批量扩展建议
- 封装下载函数:将单只股票的下载逻辑封装为函数,接收股票代码或目标URL作为参数,循环调用处理550+只股票。
- 保持会话状态:确保浏览器实例在批量任务中持续登录,避免每只股票都重新验证权限。
- 配置下载路径:通过ChromeOptions设置默认下载目录,自动保存文件无需手动选择路径。
- 添加重试机制:针对页面加载慢、按钮点击失败等情况,增加重试逻辑提升稳定性。
- 控制请求频率:添加随机等待时间(如2-5秒),避免请求过于频繁触发网站反爬限制。
内容的提问来源于stack exchange,提问作者Learner
相关产品推荐
相关产品推荐

