如何在Python Selenium中精准定位含货币金额的最子元素?
解决方法
方法1:优化XPath定位最内层元素
通过XPath筛选出自身包含目标金额文本,但所有子元素都不包含该文本的节点,也就是最内层的元素。代码调整如下:
import re from selenium import webdriver driver = webdriver.Chrome() driver.get("你的页面URL") htmlsource = driver.page_source currencypattern = re.compile("(?<=$)\d{1,5}(?:\,\d{3})?(?:\.\d+)?") for currency_match in currencypattern.finditer(htmlsource): amount = f"${currency_match.group()}" # 构造定位最内层元素的XPath xpath = f'//*[contains(normalize-space(), "{amount}") and not(.//*[contains(normalize-space(), "{amount}")])]' elements = driver.find_elements("xpath", xpath) for elem in elements: print(f"金额: {amount}, 标签名: {elem.tag_name}") # 获取指定CSS属性,可按需扩展 css_properties = { "color": elem.value_of_css_property("color"), "font-size": elem.value_of_css_property("font-size"), "font-weight": elem.value_of_css_property("font-weight"), "text-decoration": elem.value_of_css_property("text-decoration") } print("CSS属性:", css_properties)
方法2:直接用XPath正则匹配(无需先解析源码)
如果浏览器支持XPath 2.0(Chrome、Firefox等主流浏览器均支持),可以直接用XPath的matches函数匹配货币格式,同时定位最内层元素,跳过源码正则解析步骤:
from selenium import webdriver driver = webdriver.Chrome() driver.get("你的页面URL") # XPath匹配$开头的金额格式,同时确保是最内层元素 xpath = '//*[matches(normalize-space(.), "^\\$\\d{1,5}(?:,\\d{3})?(?:\\.\\d+)?$") and not(.//*[matches(normalize-space(.), "^\\$\\d{1,5}(?:,\\d{3})?(?:\\.\\d+)?$")])]' price_elements = driver.find_elements("xpath", xpath) for elem in price_elements: amount = elem.text.strip() print(f"金额: {amount}, 标签名: {elem.tag_name}") # 获取CSS属性示例 print("文本颜色:", elem.value_of_css_property("color")) print("字体大小:", elem.value_of_css_property("font-size"))
核心逻辑说明
两种方法都通过not(.//*[contains(...)])或not(.//*[matches(...)])排除了父节点——因为父节点的子元素已经包含目标金额文本,所以父节点会被过滤,最终只保留最内层的那个直接包含金额的元素。
内容的提问来源于stack exchange,提问作者Milinekticker1
相关产品推荐
相关产品推荐

