Selenium获取Excel文件下载URL时返回about:blank的问题
解决思路与方案
针对你遇到的Excel文件下载URL返回about:blank的问题,给出以下几个可行的解决方向:
1. 通过Chrome DevTools Protocol(CDP)拦截下载请求获取真实URL
Excel文件的下载请求通常是在主页面发起的,新打开的about:blank标签页只是浏览器触发下载的载体,真实的下载链接在网络请求里。可以用CDP监听网络请求,直接捕获下载URL:
代码示例:
# 初始化ChromeDriver时启用Network域 from selenium import webdriver from selenium.webdriver.chrome.options import Options options = Options() # 配置下载路径,避免弹窗 prefs = { "download.default_directory": "/your/download/path", "download.prompt_for_download": False } options.add_experimental_option("prefs", prefs) driver = webdriver.Chrome(options=options) driver.execute_cdp_cmd("Network.enable", {}) download_urls = [] # 定义请求监听回调 def capture_download_request(request): # 根据实际情况调整过滤条件,比如包含.xlsx、download等关键字 request_url = request["request"]["url"] if ".xlsx" in request_url or "download" in request_url.lower(): download_urls.append(request_url) # 添加监听事件 driver.execute_cdp_cmd("Network.requestWillBeSent", {}, capture_download_request) # 执行你的下载操作(点击Excel报告按钮) # ... 这里是你原有的点击代码 ... # 操作完成后从download_urls中获取真实链接 for url in download_urls: print_url.append(url)
2. 直接提取下载按钮的链接属性
很多网站的下载按钮会把真实链接存在href、data-href或data-url等属性里,不需要点击打开新标签页,直接读取这些属性即可:
代码示例:
reports = ["report1_xpath", "report2_xpath", "report3_xpath", "report4_xpath"] for report in reports: report_elem = self.driver.find_element(By.XPATH, report) # 先尝试获取href属性 real_url = report_elem.get_attribute("href") if not real_url or "about:blank" in real_url: # 尝试其他可能的属性,比如data-url、data-download-link等 real_url = report_elem.get_attribute("data-url") if real_url: print_url.append(real_url) # 如果需要触发下载,再点击按钮 report_elem.click()
3. 优化窗口切换与等待逻辑
原代码中固定sleep可能导致等待不充分,且切换窗口的逻辑有问题(i变量未定义)。可以用显式等待确保新窗口加载完成,再获取URL:
代码示例:
from selenium.webdriver.support.ui import WebDriverWait from selenium.common.exceptions import TimeoutException reports = ["report1_xpath", "report2_xpath", "report3_xpath", "report4_xpath"] for report in reports: original_windows = self.driver.window_handles self.driver.find_element(By.XPATH, report).click() # 等待新窗口出现 WebDriverWait(self.driver, 10).until(lambda d: len(d.window_handles) > len(original_windows)) # 获取新窗口句柄 new_window = [w for w in self.driver.window_handles if w not in original_windows][0] self.driver.switch_to.window(new_window) # 等待URL不再是about:blank try: WebDriverWait(self.driver, 15).until(lambda d: d.current_url != "about:blank") current_url = self.driver.current_url print_url.append(current_url) except TimeoutException: # 如果超时,说明是直接下载,改用CDP方法获取 pass # 关闭新窗口并切回原窗口 self.driver.close() self.driver.switch_to.window(original_windows[0])
4. 调整Chrome下载配置
通过ChromeOptions设置,让Excel直接下载而不打开新标签页,避免切换窗口的麻烦:
代码示例:
prefs = { "download.default_directory": "/your/download/path", "download.prompt_for_download": False, "download.directory_upgrade": True, # 禁用Excel预览,直接下载 "plugins.plugins_disabled": ["Microsoft Excel Viewer"] } options = Options() options.add_experimental_option("prefs", prefs) driver = webdriver.Chrome(options=options)
内容的提问来源于stack exchange,提问作者Vanhelssy
相关产品推荐
相关产品推荐

