You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium获取页面所有H2元素的文本与HREF?现有代码仅返回首个元素

解决Selenium批量获取H2标签文本及对应链接的问题

你当前代码只返回首个元素,是因为使用的visibility_of_element_located仅定位单个匹配元素。要获取所有符合条件的H2标签,需要调整等待条件为获取多个元素的方法,同时遍历元素提取所需内容。

修改后的代码示例

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

options = webdriver.ChromeOptions()
driver = webdriver.Chrome(options=options)
driver.maximize_window()
wait = WebDriverWait(driver, 10)

url = 'http://www.biblioteca.presidencia.gov.br/presidencia/ex-presidentes/jose-sarney/discursos/1985?b_start:int=0'
driver.get(url)

# 等待所有class为tileHeadline的H2元素加载完成
h2_elements = wait.until(EC.visibility_of_all_elements_located((By.CLASS_NAME, "tileHeadline")))

# 遍历每个H2元素,提取文本和对应链接
for h2 in h2_elements:
    # 获取H2的文本内容
    h2_text = h2.text
    # 获取H2内部a标签的href属性
    h2_href = h2.find_element(By.TAG_NAME, "a").get_attribute("href")
    print(f"文本: {h2_text}")
    print(f"链接: {h2_href}")
    print("---")

driver.quit()

关键说明

  • 用visibility_of_all_elements_located替代visibility_of_element_located,确保等待所有目标元素加载可见后再返回元素列表。
  • 目标页面的H2标签内,链接通常嵌套在<a>子标签中,通过find_element(By.TAG_NAME, "a")定位子元素后,用get_attribute("href")即可提取完整链接地址。
  • 若需要提取<a>标签的文本而非H2的文本,可直接替换为h2.find_element(By.TAG_NAME, "a").text,可根据页面实际结构灵活调整。

内容的提问来源于stack exchange,提问作者Marcelo Soares

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 16:03:23