Selenium技术问题:获取URL后requests无法下载及无头模式下载路径失效
Hey there! Let's work through these two tricky Selenium problems you're facing—they're super common, so you're not alone.
1. Fixing Requests Download Failures After Getting URLs via Selenium
The most likely culprit here is that the URL you're grabbing requires authenticated session data (like cookies) or specific request headers that Requests isn't sending. Selenium runs in a browser that's already logged in/has the necessary session context, but Requests starts with a fresh, unauthenticated session.
How to Fix It:
Pass the browser's cookies and matching headers from Selenium to Requests to replicate the authenticated session:
from selenium import webdriver from selenium.webdriver.common.by import By import requests # Step 1: Get the URL and session data with Selenium driver = webdriver.Chrome() driver.get("your_target_page_url") # Example: Extract the file download URL from the page file_url = driver.find_element(By.XPATH, "//a[@download]").get_attribute("href") # Step 2: Extract cookies from Selenium and add to a Requests session session = requests.Session() for cookie in driver.get_cookies(): session.cookies.set(cookie['name'], cookie['value']) # Step 3: Copy critical headers from the browser to avoid anti-scraping blocks headers = { "User-Agent": driver.execute_script("return navigator.userAgent;"), "Referer": driver.current_url # Some sites check referrer to prevent hotlinking } # Step 4: Download the file with the authenticated session response = session.get(file_url, headers=headers, stream=True) with open("downloaded_file.pdf", "wb") as f: for chunk in response.iter_content(chunk_size=8192): f.write(chunk) driver.quit()
Why This Works:
- The session carries over login cookies from Selenium, so the server recognizes you as an authenticated user.
- Matching the User-Agent and Referer headers avoids triggering anti-scraping checks that block Requests' default generic headers.
2. Fixing Download Path Reset in Headless Mode
Older versions of Chrome's headless mode had limitations that ignored custom download preferences. The fix uses the new headless mode (--headless=new) and explicitly configures download settings in the browser options.
How to Fix It:
Configure your ChromeOptions with both headless mode and persistent download preferences:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By # Set up Chrome options chrome_options = Options() # Use the NEW headless mode (critical for download settings to work properly) chrome_options.add_argument("--headless=new") # Disable GPU (optional but recommended for headless environments) chrome_options.add_argument("--disable-gpu") # Define your custom download path (use absolute paths!) download_path = "/path/to/your/custom/download/folder" # On Windows, use raw strings or escaped slashes: r"C:\Users\YourName\Downloads" prefs = { "download.default_directory": download_path, "download.prompt_for_download": False, # Disable download prompt "profile.default_content_settings.popups": 0, # Block popups that interfere with downloads "safebrowsing.enabled": True # Optional: Enable safe browsing if needed } chrome_options.add_experimental_option("prefs", prefs) # Initialize driver with the configured options driver = webdriver.Chrome(options=chrome_options) # Test the download workflow driver.get("page_with_download_button") driver.find_element(By.ID, "download-button").click() # Add explicit waits here if needed to let the download finish driver.quit()
Key Notes:
- Use
--headless=newinstead of the old--headless: The original headless mode ran a stripped-down browser that didn't respect many preferences. The new mode behaves exactly like regular Chrome. - Always use an absolute path for the download folder—relative paths can cause unexpected behavior across different systems.
内容的提问来源于stack exchange,提问作者Aniket

