Selenium find_element定位返回部分网页内容问题排查
爬取HHI品种保证金表格的元素定位问题
问题描述
- 完成Python自动化爬虫教程学习后,尝试修改代码爬取目标站点的HHI品种保证金表格,因目标网站代码结构特殊,元素定位存在较大障碍。
- 已通过Xpath表达式
//a[@name="HHI"]定位到HHI对应的锚点子元素,该元素的父节点为<font size="2"></font>标签,标签内部包含需要提取的保证金表格文本;但页面中存在大量属性完全相同的<font size="2"></font>标签,无法直接通过Xpath//font[@size="2"]完成精准定位。 - 尝试使用全路径绝对Xpath定位时,返回结果包含了近半网页的冗余内容,无法精准提取目标文本,所用的绝对Xpath为多层嵌套font标签的超长路径:
/html/body/table/tbody/tr/td/table/tbody/tr/td/table/tbody/tr[3]/td/pre/font/table/tbody/tr/td[2]/pre/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font/font
- 参考教程为freeCodeCamp发布的Python自动化全入门课程。
初始实现代码
from selenium.webdriver.chrome.options import Options from selenium.webdriver.chrome.service import Service import pandas as pd # prepare it to automate from datetime import datetime import os import sys import csv application_path = os.path.dirname(sys.executable) # export the result to the same file as the executable now = datetime.now() # for modify the export name with a date month_day_year = now.strftime("%m%d%Y") # MMDDYYYY website = "https://www.hkex.com.hk/eng/market/rm/rm_dcrm/riskdata/margin_hkcc/merte_hkcc.htm" path = "C:/Users/User/PycharmProjects/Automate with Python – Full Course for Beginners/venv/Scripts/chromedriver.exe" # headless-mode options = Options() options.headless = True service = Service(executable_path=path) driver = webdriver.Chrome(service=service, options=options) driver.get(website) containers = driver.find_element(by="xpath", value='') # or find_elements hhi = containers.text # if using find_elements, = containers[0].text print(hhi)
临时可行方案
- 经Xpath语法调整后,即使定位到准确的
font标签,受页面嵌套标签不规范的结构影响,返回结果仍会包含后续所有标签的全部文本,无法直接拿到单一品种的内容。 - 目前验证可正常运行的处理逻辑:
- 用Xpath
//font[a/@name="{product}"]定位到对应品种锚点所在的font标签 - 调用
.split("Back to Top")方法按页面固定的返回顶部标识拆分不同产品的内容,生成列表后取首项,即可得到HHI品种的独立文本块 - 调用
.split("\n")按换行符拆分文本,后续可进一步处理嵌套列表,最终整理为以行权价为索引、到期日为列名的pandas DataFrame结构
- 用Xpath
- 该方案虽然执行效率不是最高,但目前可稳定运行,调整后的实现代码如下:
product = "HHI" containers = driver.find_element(by="xpath", value=f'//font[a/@name="{product}"]') hhi = containers.text.split("Back to Top") # print(hhi) hhi1 = hhi[0].split("\n") df = pd.DataFrame(hhi1) # print(df) df.to_csv(f"{product}_{month_day_year}.csv")
内容的提问来源于stack exchange,提问作者Stephen
相关产品推荐
相关产品推荐

