如何使用Selenium下载新标签页中打开的文件?OSC网站文件下载遇TimeoutException问题求助
Troubleshooting TimeoutException & Tab Switching Issues in Selenium for OSC Bulletin Downloads
Let's break down why you're hitting that TimeoutException and fix the tab switching + element locating problems step by step:
Common Issues in Your Current Code
- Unreliable Tab Switching: You're switching to
driver.window_handles[1]immediately after opening the link, but the new tab might not have finished loading yet. This can lead to trying to interact with elements that don't exist in the active tab. - Fragile XPATH: That super-long absolute XPATH is prone to breaking if the page structure changes even slightly. Relative locators are much more stable.
- Incorrect Download Target: You're targeting the
<svg>element directly, but often the clickable trigger is its parent element (like an<a>or<button>) rather than the icon itself.
Fixed Code with Explanations
Here's a revised version of your script that addresses these issues:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.common.keys import Keys from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException import time # Initialize driver service = webdriver.ChromeService(r"C:\Users\Lenovo\Desktop\chromedriver.exe") driver = webdriver.Chrome(service=service) driver.maximize_window() try: # Navigate to the search results directly (saves a step instead of searching manually) driver.get('https://www.osc.ca/en/securities-law/osc-bulletin?keyword=61-101&date%5Bmin%5D=&date%5Bmax%5D=&sort_bef_combine=field_start_date_DESC') # Wait for search results to load WebDriverWait(driver, 15).until( EC.presence_of_element_located((By.CSS_SELECTOR, "div.view-content div.media-body a")) ) # Get all result links (limit to first 20 as in your original code) result_links = driver.find_elements(By.CSS_SELECTOR, "div.view-content div.media-body a")[:20] for link in result_links: # Store original tab handle before opening new tab original_tab = driver.current_window_handle # Open link in new tab (using Ctrl+Enter to avoid unexpected navigation) link.send_keys(Keys.CONTROL + Keys.ENTER) # Wait for new tab to open and switch to it WebDriverWait(driver, 10).until(lambda d: len(d.window_handles) > 1) new_tab = [tab for tab in driver.window_handles if tab != original_tab][0] driver.switch_to.window(new_tab) try: # Wait for the download button (target the parent element, not the SVG) # Adjust selector if needed - inspect the page to find the clickable parent download_btn = WebDriverWait(driver, 15).until( EC.element_to_be_clickable((By.XPATH, "//a[contains(@href, '.pdf')]")) ) download_btn.click() # Optional: Wait for download to start (adjust time based on your connection) time.sleep(3) except TimeoutException: print(f"Could not find download button for tab: {new_tab}") finally: # Close new tab and switch back to original driver.close() driver.switch_to.window(original_tab) finally: driver.quit()
Key Improvements
- Direct Search URL: Instead of manually typing the search term, we navigate straight to the pre-filtered results page - this eliminates potential issues with the search form.
- Stable Locators: We use CSS selectors and relative XPATHs that are less likely to break if the page layout changes.
- Safe Tab Switching: We wait for the new tab to exist before switching, and explicitly identify the new tab handle instead of relying on index
[1]. - Error Handling: Added try/finally blocks to ensure tabs are closed and we switch back to the original tab even if a download fails.
- Target Correct Download Element: We look for a link with a PDF href (common for download buttons) instead of the SVG icon, which is usually just a visual element.
Additional Tips
- If the download button still isn't found, right-click the button on the page and select "Inspect" to find its actual HTML structure. Adjust the locator to match the clickable element (often an
<a>with a download attribute or a<button>). - Avoid using
time.sleep()where possible, but a short sleep after clicking download can help ensure the file starts downloading before moving to the next item. - Make sure your ChromeDriver version matches your installed Chrome browser version - mismatches can cause unexpected behavior.
内容的提问来源于stack exchange,提问作者Nuri Taş
相关产品推荐
相关产品推荐

