You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Selenium下载新标签页中打开的文件?OSC网站文件下载遇TimeoutException问题求助

Troubleshooting TimeoutException & Tab Switching Issues in Selenium for OSC Bulletin Downloads

Let's break down why you're hitting that TimeoutException and fix the tab switching + element locating problems step by step:

Common Issues in Your Current Code

  • Unreliable Tab Switching: You're switching to driver.window_handles[1] immediately after opening the link, but the new tab might not have finished loading yet. This can lead to trying to interact with elements that don't exist in the active tab.
  • Fragile XPATH: That super-long absolute XPATH is prone to breaking if the page structure changes even slightly. Relative locators are much more stable.
  • Incorrect Download Target: You're targeting the <svg> element directly, but often the clickable trigger is its parent element (like an <a> or <button>) rather than the icon itself.

Fixed Code with Explanations

Here's a revised version of your script that addresses these issues:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.common.keys import Keys
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException
import time

# Initialize driver
service = webdriver.ChromeService(r"C:\Users\Lenovo\Desktop\chromedriver.exe")
driver = webdriver.Chrome(service=service)
driver.maximize_window()

try:
    # Navigate to the search results directly (saves a step instead of searching manually)
    driver.get('https://www.osc.ca/en/securities-law/osc-bulletin?keyword=61-101&date%5Bmin%5D=&date%5Bmax%5D=&sort_bef_combine=field_start_date_DESC')
    
    # Wait for search results to load
    WebDriverWait(driver, 15).until(
        EC.presence_of_element_located((By.CSS_SELECTOR, "div.view-content div.media-body a"))
    )
    
    # Get all result links (limit to first 20 as in your original code)
    result_links = driver.find_elements(By.CSS_SELECTOR, "div.view-content div.media-body a")[:20]
    
    for link in result_links:
        # Store original tab handle before opening new tab
        original_tab = driver.current_window_handle
        
        # Open link in new tab (using Ctrl+Enter to avoid unexpected navigation)
        link.send_keys(Keys.CONTROL + Keys.ENTER)
        
        # Wait for new tab to open and switch to it
        WebDriverWait(driver, 10).until(lambda d: len(d.window_handles) > 1)
        new_tab = [tab for tab in driver.window_handles if tab != original_tab][0]
        driver.switch_to.window(new_tab)
        
        try:
            # Wait for the download button (target the parent element, not the SVG)
            # Adjust selector if needed - inspect the page to find the clickable parent
            download_btn = WebDriverWait(driver, 15).until(
                EC.element_to_be_clickable((By.XPATH, "//a[contains(@href, '.pdf')]"))
            )
            download_btn.click()
            
            # Optional: Wait for download to start (adjust time based on your connection)
            time.sleep(3)
            
        except TimeoutException:
            print(f"Could not find download button for tab: {new_tab}")
        finally:
            # Close new tab and switch back to original
            driver.close()
            driver.switch_to.window(original_tab)
            
finally:
    driver.quit()

Key Improvements

  • Direct Search URL: Instead of manually typing the search term, we navigate straight to the pre-filtered results page - this eliminates potential issues with the search form.
  • Stable Locators: We use CSS selectors and relative XPATHs that are less likely to break if the page layout changes.
  • Safe Tab Switching: We wait for the new tab to exist before switching, and explicitly identify the new tab handle instead of relying on index [1].
  • Error Handling: Added try/finally blocks to ensure tabs are closed and we switch back to the original tab even if a download fails.
  • Target Correct Download Element: We look for a link with a PDF href (common for download buttons) instead of the SVG icon, which is usually just a visual element.

Additional Tips

  • If the download button still isn't found, right-click the button on the page and select "Inspect" to find its actual HTML structure. Adjust the locator to match the clickable element (often an <a> with a download attribute or a <button>).
  • Avoid using time.sleep() where possible, but a short sleep after clicking download can help ensure the file starts downloading before moving to the next item.
  • Make sure your ChromeDriver version matches your installed Chrome browser version - mismatches can cause unexpected behavior.

内容的提问来源于stack exchange,提问作者Nuri Taş

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 00:49:08