如何通过Selenium与XPath提取HTML中独立文本形式的物料计量单位?
解决方法:提取div中的独立文本节点
首先,你遇到的问题根源很明确:
- 你获取的是整个div的
innerText,里面混杂了价格和目标文本/each,而你的正则表达式^[a-zA-Z]*$要求从开头到结尾全是字母,自然匹配不到有效内容,返回None,导致后续拼接字符串时报错。 - 目标文本
/each是div元素的直接文本节点,不属于任何子元素,所以需要精准定位这个文本节点来提取。
正确的XPath方案
这里有两种简洁有效的方式可以提取到each:
方案1:定位直接文本节点后清理内容
先用XPath选中div下的非空白直接文本节点,再手动去除多余的符号和空格:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.common.exceptions import NoSuchElementException try: # 定位div下的非空白直接文本节点 text_element = driver.find_element(By.XPATH, "//div[@class='pricingReg']/text()[normalize-space()]") # 获取文本内容并清理格式 raw_text = text_element.get_attribute('textContent').strip() purchase_unit_label = raw_text.lstrip('/').strip() # 去掉前面的斜杠和空格 print(f'Purchase Unit Label = {purchase_unit_label}') except NoSuchElementException: print('Purchase Unit Label NoSuchElementException')
方案2:用XPath函数直接提取目标文本
利用XPath的字符串处理能力,一步到位截取到each:
try: # 直接通过XPath提取斜杠后的目标文本 purchase_unit_label = driver.find_element(By.XPATH, "substring-after(normalize-space(//div[@class='pricingReg']/text()[normalize-space()]), '/')" ).get_attribute('textContent').strip() print(f'Purchase Unit Label = {purchase_unit_label}') except NoSuchElementException: print('Purchase Unit Label NoSuchElementException')
方法原理说明
//div[@class='pricingReg']/text():专门选取div元素下的所有直接文本节点(不会包含子元素的文本内容)。[normalize-space()]:过滤掉空的文本节点,只保留有实际内容的节点(这里就是/each)。substring-after(..., '/'):直接截取斜杠后面的部分,省去手动处理的步骤。
内容的提问来源于stack exchange,提问作者Kyle Linden
相关产品推荐
相关产品推荐

