如何用Python Selenium从世界银行网站下载无直接链接的Excel文件?
用Selenium自动化下载世界银行制裁企业Excel文件
我尝试从世界银行制裁企业页面下载Excel文件,想用Python的Selenium实现自动化下载,但试了XPath、类选择器和CSS选择器都没成功。以下是我的代码,求技术建议:
from selenium import webdriver from selenium.webdriver.common.keys import Keys from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.common.by import By from selenium.webdriver.chrome.service import Service from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC options = webdriver.ChromeOptions() options.add_argument("--log-level=OFF") options.add_experimental_option('excludeSwitches', ['enable-logging']) driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options) try: #driver.get('https://www.worldbank.org/en/projects-operations/procurement/debarred-firms'); #downloadcsv= driver.find_element(By.XPATH, '//*[@id="k-debarred-firms"]/div[1]/a'); #gotit= driver.find_element(By.CLASS_NAME, "dialog_form_actions"); #gotit.click(); #Click on Download Button #driver.find_element(By.XPATH,'//*[@id="k-debarred-firms"]/div[1]/a').click() #time.sleep(50) driver.execute("get", {'url': 'https://www.worldbank.org/en/projects-operations/procurement/debarred-firms'}) WebDriverWait(driver, 200).until(EC.element_to_be_clickable((By.CSS_SELECTOR, 'title="Excel"'))).click() driver.close() print("file downloaded") except: print("Invalid URL")
问题分析与修正方案
原代码核心问题
- CSS选择器语法错误:
title="Excel"不是合法的CSS选择器,正确写法应为a[title="Excel"],用来定位带title="Excel"属性的<a>标签 - 未处理页面弹窗:页面加载后会出现隐私提示弹窗,必须先关闭才能操作下载按钮
- 异常捕获过于宽泛:直接用
except:会掩盖所有错误,无法定位具体问题 - 页面跳转写法冗余:
driver.execute("get")不如直接用driver.get(url)简洁直观
修正后的代码
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from webdriver_manager.chrome import ChromeDriverManager import time # 配置Chrome选项,可指定默认下载路径(避免浏览器下载确认弹窗) options = webdriver.ChromeOptions() options.add_argument("--log-level=OFF") options.add_experimental_option('excludeSwitches', ['enable-logging']) # 替换成你的本地下载路径 prefs = {"download.default_directory": "C:/Your/Download/Path"} options.add_experimental_option("prefs", prefs) driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options) try: # 打开目标页面 driver.get('https://www.worldbank.org/en/projects-operations/procurement/debarred-firms') # 等待并关闭隐私提示弹窗 WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.CSS_SELECTOR, "button.dialog_form_actions")) ).click() # 等待Excel下载按钮可点击并触发点击 excel_download_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.CSS_SELECTOR, 'a[title="Excel"]')) ) excel_download_btn.click() # 等待文件下载完成(根据文件大小调整时长) time.sleep(10) print("文件下载完成") except Exception as e: print(f"出错了: {str(e)}") finally: # 确保浏览器关闭 driver.quit()
关键说明
- 弹窗处理:用
button.dialog_form_actions定位隐私提示的确认按钮,消除后续操作的干扰 - 精准元素定位:通过
a[title="Excel"]定位下载按钮,比XPath更稳定 - 下载路径配置:设置
download.default_directory可指定文件保存位置,避免手动确认 - 异常处理优化:捕获具体异常并打印信息,方便排查问题
- 显式等待替代固定延时:用
WebDriverWait等待元素加载,避免因页面加载慢导致的元素未找到问题
内容的提问来源于stack exchange,提问作者Raja Manikam
相关产品推荐
相关产品推荐

