You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Selenium与Scrapy点击页面标题链接抓取数据:代码错误排查及实现方法咨询

First, let's quickly touch on why your original code hit a snag: the XPath //tbody//td[6] might have been targeting an element that doesn’t exist, is hidden, or there could be multiple tbody elements on the page throwing off the selector. But since you’ve shifted focus to clicking title links and scraping their detail pages, let’s walk through a robust, practical implementation for that.

Step-by-Step Implementation

Here’s a fully commented solution that handles navigating to the main page, clicking title links, scraping detail page data, and returning to the main list—with safeguards against common Selenium pitfalls:

from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
import time

# Initialize Chrome driver (double-check your chromedriver path matches your system)
driver = webdriver.Chrome('C:\\Program Files (x86)\\chromedriver.exe')
main_url = 'https://www.aeafa.es/asociados.php?provinput='
driver.get(main_url)
wait = WebDriverWait(driver, 30)

try:
    # Wait for the main table to load (adjust selector if your table has a specific ID/class)
    wait.until(EC.presence_of_element_located((By.TAG_NAME, "table")))
    
    # Fetch all title links (update this XPath to match your actual title elements!)
    # Example: if titles are <a> tags inside table cells with a "detalle" href, use this:
    title_links = wait.until(EC.presence_of_all_elements_located(
        (By.XPATH, "//tbody//td//a[contains(@href, 'detalle')]")
    ))
    
    # Iterate through each link (re-fetch links each loop to avoid stale element errors)
    for i in range(len(title_links)):
        # Re-grab the links list to prevent StaleElementReferenceException
        title_links = wait.until(EC.presence_of_all_elements_located(
            (By.XPATH, "//tbody//td//a[contains(@href, 'detalle')]")
        ))
        current_link = title_links[i]
        
        # Click the title link
        current_link.click()
        
        # Wait for the detail page to fully load (use a unique element on your detail page here)
        wait.until(EC.presence_of_element_located((By.TAG_NAME, "h1")))
        
        # Scrape your target data (customize this to match what you need!)
        detail_title = driver.find_element(By.TAG_NAME, "h1").text
        contact_info = driver.find_element(By.CLASS_NAME, "contacto").text  # Example class name
        
        print(f"=== Detail for {detail_title} ===")
        print(f"Contact Info: {contact_info}\n")
        
        # Navigate back to the main page
        if len(driver.window_handles) > 1:
            # If the link opened a new tab: close the detail tab and switch back
            driver.close()
            driver.switch_to.window(driver.window_handles[0])
        else:
            # If it opened in the same tab: use back()
            driver.back()
        
        time.sleep(1)  # Short pause to avoid overwhelming the server

finally:
    # Clean up: always close the driver when done
    driver.quit()

Key Customization Tips

  • Update Selectors: Use your browser’s dev tools to inspect the title links and detail page elements, then adjust the XPath/CSS selectors in the code to match. For example, if titles have a specific class like titulo-enlace, use By.CLASS_NAME instead of XPath.
  • Detail Page Scraping: Modify the scraping section to target the exact data you need—like addresses, phone numbers, or descriptions—using the appropriate selectors (By.ID, By.CSS_SELECTOR, etc.).
  • Handle Navigation: If links open in new tabs, the window handle logic will work; if they open in the same tab, driver.back() is sufficient.
  • Avoid Stale Elements: Re-fetching the title links in each loop prevents errors caused by the page DOM changing after navigation.

Quick Troubleshooting

  • If links aren’t found: Verify your selector in the browser’s dev tools (use the "Inspect" feature to test XPath/CSS queries).
  • If elements are unclickable: Check for overlays (like modals or popups) and add waits for those to disappear before clicking.
  • Prefer explicit waits (WebDriverWait) over time.sleep() wherever possible—they make your code more reliable by waiting only as long as needed.

内容的提问来源于stack exchange,提问作者Amen Aziz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 23:59:08