使用Python和WebDriver时Selenium无法获取维基百科指定元素文本的问题求助
解决Selenium获取维基百科文章数文本报错的问题
我看了你的代码,问题出在元素引用过时(StaleElementReferenceException),原因很简单:你先定位了articles元素,然后点击了这个元素跳转到新页面,这时候原来页面的DOM已经被销毁了,再去访问之前保存的articles元素自然会报错。
另外,你用了绝对XPATH定位,这种方式非常脆弱,页面结构稍微调整就会失效,推荐用更稳定的定位方式,比如利用目标元素所在的articlecount这个id来定位。
修正后的代码
from selenium import webdriver from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC chrome_driver_path = "/Users/greta/Development/chromedriver" driver = webdriver.Chrome(executable_path=chrome_driver_path) driver.get("https://en.wikipedia.org/wiki/Main_Page") try: # 用更稳定的CSS选择器定位:通过id为articlecount的div内部的第一个a标签 article_count_element = WebDriverWait(driver, 20).until( EC.visibility_of_element_located((By.CSS_SELECTOR, "#articlecount a")) ) # 先获取文本,再执行点击(如果需要跳转的话) number_of_articles = article_count_element.text print(number_of_articles) # 如果需要点击跳转,放在获取文本之后 article_count_element.click() except Exception as e: print(f"出错了: {e}") finally: driver.quit()
关键说明
- 调整操作顺序:必须在页面跳转前获取元素的文本内容,否则原元素的引用会因为页面刷新/跳转而失效。
- 优化定位策略:
- CSS选择器
#articlecount a:直接通过id定位父div,再选中内部的a标签,比绝对XPATH可靠得多 - 也可以用相对XPATH:
//div[@id='articlecount']/a[1],同样比绝对路径稳定
- CSS选择器
- 显式等待的正确使用:用
visibility_of_element_located确保元素可见再操作,比element_to_be_clickable更适合获取文本的场景(当然如果要点击的话,后者也没问题)
这样修改后,就能正常获取到目标元素的文本内容,同时避免了元素引用过时的问题。
内容的提问来源于stack exchange,提问作者Greta
相关产品推荐
相关产品推荐

