You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium find_element定位返回部分网页内容问题排查

爬取HHI品种保证金表格的元素定位问题

问题描述

  • 完成Python自动化爬虫教程学习后,尝试修改代码爬取目标站点的HHI品种保证金表格,因目标网站代码结构特殊,元素定位存在较大障碍。
  • 已通过Xpath表达式//a[@name="HHI"]定位到HHI对应的锚点子元素,该元素的父节点为<font size="2"></font>标签,标签内部包含需要提取的保证金表格文本;但页面中存在大量属性完全相同的<font size="2"></font>标签,无法直接通过Xpath//font[@size="2"]完成精准定位。
  • 尝试使用全路径绝对Xpath定位时,返回结果包含了近半网页的冗余内容,无法精准提取目标文本,所用的绝对Xpath为多层嵌套font标签的超长路径:
/html/body/table/tbody/tr/td/table/tbody/tr/td/table/tbody/tr[3]/td/pre/font/table/tbody/tr/td[2]/pre/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font
  • 参考教程为freeCodeCamp发布的Python自动化全入门课程。

初始实现代码

from selenium.webdriver.chrome.options import Options
from selenium.webdriver.chrome.service import Service
import pandas as pd
# prepare it to automate
from datetime import datetime
import os
import sys
import csv

application_path = os.path.dirname(sys.executable) # export the result to the same file as the executable

now = datetime.now() # for modify the export name with a date
month_day_year = now.strftime("%m%d%Y") # MMDDYYYY

website = "https://www.hkex.com.hk/eng/market/rm/rm_dcrm/riskdata/margin_hkcc/merte_hkcc.htm"
path = "C:/Users/User/PycharmProjects/Automate with Python – Full Course for Beginners/venv/Scripts/chromedriver.exe"

# headless-mode
options = Options()
options.headless = True

service = Service(executable_path=path)
driver = webdriver.Chrome(service=service, options=options)
driver.get(website)

containers = driver.find_element(by="xpath", value='') # or find_elements

hhi = containers.text # if using find_elements, = containers[0].text

print(hhi)

临时可行方案

  • 经Xpath语法调整后,即使定位到准确的font标签,受页面嵌套标签不规范的结构影响,返回结果仍会包含后续所有标签的全部文本,无法直接拿到单一品种的内容。
  • 目前验证可正常运行的处理逻辑:
    • 用Xpath//font[a/@name="{product}"]定位到对应品种锚点所在的font标签
    • 调用.split("Back to Top")方法按页面固定的返回顶部标识拆分不同产品的内容,生成列表后取首项,即可得到HHI品种的独立文本块
    • 调用.split("\n")按换行符拆分文本,后续可进一步处理嵌套列表,最终整理为以行权价为索引、到期日为列名的pandas DataFrame结构
  • 该方案虽然执行效率不是最高,但目前可稳定运行,调整后的实现代码如下:
product = "HHI"

containers = driver.find_element(by="xpath", value=f'//font[a/@name="{product}"]')

hhi = containers.text.split("Back to Top")

# print(hhi)

hhi1 = hhi[0].split("\n")

df = pd.DataFrame(hhi1)

# print(df)

df.to_csv(f"{product}_{month_day_year}.csv")

内容的提问来源于stack exchange,提问作者Stephen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 04:45:37