You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python和WebDriver时Selenium无法获取维基百科指定元素文本的问题求助

解决Selenium获取维基百科文章数文本报错的问题

我看了你的代码,问题出在元素引用过时(StaleElementReferenceException),原因很简单:你先定位了articles元素,然后点击了这个元素跳转到新页面,这时候原来页面的DOM已经被销毁了,再去访问之前保存的articles元素自然会报错。

另外,你用了绝对XPATH定位,这种方式非常脆弱,页面结构稍微调整就会失效,推荐用更稳定的定位方式,比如利用目标元素所在的articlecount这个id来定位。

修正后的代码

from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC

chrome_driver_path = "/Users/greta/Development/chromedriver"
driver = webdriver.Chrome(executable_path=chrome_driver_path)
driver.get("https://en.wikipedia.org/wiki/Main_Page")

try:
    # 用更稳定的CSS选择器定位:通过id为articlecount的div内部的第一个a标签
    article_count_element = WebDriverWait(driver, 20).until(
        EC.visibility_of_element_located((By.CSS_SELECTOR, "#articlecount a"))
    )
    # 先获取文本,再执行点击(如果需要跳转的话)
    number_of_articles = article_count_element.text
    print(number_of_articles)
    
    # 如果需要点击跳转,放在获取文本之后
    article_count_element.click()
except Exception as e:
    print(f"出错了: {e}")
finally:
    driver.quit()

关键说明

  1. 调整操作顺序:必须在页面跳转前获取元素的文本内容,否则原元素的引用会因为页面刷新/跳转而失效。
  2. 优化定位策略:
    • CSS选择器#articlecount a:直接通过id定位父div,再选中内部的a标签,比绝对XPATH可靠得多
    • 也可以用相对XPATH://div[@id='articlecount']/a[1],同样比绝对路径稳定
  3. 显式等待的正确使用:用visibility_of_element_located确保元素可见再操作,比element_to_be_clickable更适合获取文本的场景(当然如果要点击的话,后者也没问题)

这样修改后,就能正常获取到目标元素的文本内容,同时避免了元素引用过时的问题。

内容的提问来源于stack exchange,提问作者Greta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 17:37:45