You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium WebDriver与Python提取元素文本及指定区域内容?

解决Selenium提取网页文本的问题

嗨,我来帮你搞定这个文本提取的问题!首先得纠正你之前的一个关键错误:getWindowHandle()是用来获取浏览器窗口句柄的(比如切换多窗口时用),完全不是提取元素文本的方法,所以你调用它肯定拿不到想要的内容。要提取元素内的文本,我们得用Selenium元素对象的.text属性。

接下来针对你要提取的目标节点<span translate="no">大塊文化</span>,我给你两种可靠的实现方案:

方法1:用XPath定位元素

XPath可以精准匹配到这个translate属性为no的span标签,再通过.text提取文本:

# 定位目标span元素
publisher_element = driver.find_element_by_xpath('//span[@translate="no"]')
# 提取并打印文本
publisher_text = publisher_element.text
print(publisher_text)  # 输出:大塊文化

方法2:用CSS选择器定位元素

CSS选择器同样能精准匹配目标元素,代码更简洁:

# 定位目标span元素
publisher_element = driver.find_element_by_css_selector('span[translate="no"]')
# 提取并打印文本
publisher_text = publisher_element.text
print(publisher_text)  # 输出:大塊文化

顺便纠正你之前的书名提取代码

你之前用find_elements_by_xpath返回的是元素列表,提取文本时同样要用.text,而不是错误的getWindowHandle():

# 单个元素建议用find_element_by_xpath(无需处理列表)
book_title_element = driver.find_element_by_xpath('//p[@class="title product-field"]')
book_title = book_title_element.text
print(book_title)

实用小提示

如果页面加载较慢,记得加上显式等待避免元素未加载导致报错:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

# 最多等待10秒,直到目标元素加载完成
publisher_element = WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.XPATH, '//span[@translate="no"]'))
)
publisher_text = publisher_element.text

内容的提问来源于stack exchange,提问作者hakukou

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:02:41