如何用Python+Selenium获取悬停元素中的标签数值?
解决方法
你需要的数值45、77、298并没有直接渲染在页面DOM中,而是存储在目标元素的data-content属性里的HTML字符串中。直接调用.text只能获取页面上可见的文本(也就是420和Guru),所以得换个思路:
步骤1:获取data-content属性的内容
先定位到id="guru"的元素,然后提取它的data-content属性值:
from selenium.webdriver.common.by import By guru_element = driver.find_element(By.ID, 'guru') data_content = guru_element.get_attribute('data-content')
步骤2:解析HTML字符串提取数值
拿到data-content的内容后,用BeautifulSoup来解析这段HTML,提取所有<span>标签里的文本:
from bs4 import BeautifulSoup soup = BeautifulSoup(data_content, 'html.parser') span_texts = [span.get_text(strip=True) for span in soup.find_all('span')] print(span_texts) # 输出: ['45', '77', '298']
可选:用JavaScript直接提取(无需额外库)
如果不想引入BeautifulSoup,也可以用Selenium执行JavaScript来解析属性里的HTML:
script = """ const elem = document.getElementById('guru'); const content = elem.getAttribute('data-content'); const tempDiv = document.createElement('div'); tempDiv.innerHTML = content; const spans = tempDiv.querySelectorAll('span'); return Array.from(spans).map(s => s.textContent.trim()); """ values = driver.execute_script(script) print(values) # 输出: ['45', '77', '298']
原理说明:这些数值所在的HTML结构是作为属性值存储的,只有当鼠标悬停时,页面才会把这段HTML渲染成可见的弹窗。所以直接通过Selenium查找页面上的<span>元素是找不到的,必须先提取属性里的HTML字符串再解析。
内容的提问来源于stack exchange,提问作者Chuck
相关产品推荐
相关产品推荐

