You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium提取文本为LEIA ESTA EDIÇÃO的多个Href链接

Got it, let's get this sorted for you! The key here is to reliably locate that specific <a> tag and pull its href attribute—here's a robust implementation using Selenium, with explanations for each step:

Step-by-Step Solution

First, we'll use explicit waits (way more reliable than implicit waits) to make sure the element loads before we try to interact with it, and we'll cover multiple ways to locate the link in case one doesn't work due to rendering quirks.

Full Working Code

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException

# Initialize Chrome driver
driver = webdriver.Chrome()

try:
    # Replace this with your actual target website URL
    driver.get("https://your-target-website.com")

    # Wait up to 10 seconds for the link to appear (adjust timeout as needed)
    wait = WebDriverWait(driver, 10)

    # Option 1: Locate using exact link text (matches the visible text)
    # This works if the text is exactly "LEIA ESTA EDIÇÃO" (including the Ç)
    link_element = wait.until(EC.presence_of_element_located((By.LINK_TEXT, "LEIA ESTA EDIÇÃO")))

    # Option 2: Fallback with XPath (more flexible if text has hidden whitespace/encoding issues)
    # link_element = wait.until(EC.presence_of_element_located((By.XPATH, "//a[normalize-space(text())='LEIA ESTA EDIÇÃO']")))

    # Option 3: Use the title attribute (since your example has this exact title)
    # link_element = wait.until(EC.presence_of_element_located((By.XPATH, "//a[@title='LEIA ESTA EDIÇÃO']")))

    # Extract the href value
    target_link = link_element.get_attribute("href")
    print(f"Successfully extracted link: {target_link}")

except TimeoutException:
    print("Oops! The 'LEIA ESTA EDIÇÃO' link didn't load within 10 seconds, or wasn't found on the page.")

finally:
    # Always clean up by closing the driver
    driver.quit()

Key Explanations

  • Explicit Waits: Using WebDriverWait ensures we don't try to find the element before it's rendered on the page—this fixes most "element not found" errors.
  • Multiple Locator Options:
    • By.LINK_TEXT: Directly matches the visible text of the link (best if the text is consistent).
    • XPath with normalize-space: Handles cases where the text has hidden spaces or line breaks.
    • XPath with @title: Uses the title attribute from your example, which is often a reliable fallback if the visible text changes.
  • Error Handling: The TimeoutException catch lets you gracefully handle cases where the link isn't present or takes too long to load.

内容的提问来源于stack exchange,提问作者Luís Henrique Martins

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:42:31