You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python Selenium从世界银行网站下载无直接链接的Excel文件?

用Selenium自动化下载世界银行制裁企业Excel文件

我尝试从世界银行制裁企业页面下载Excel文件,想用Python的Selenium实现自动化下载,但试了XPath、类选择器和CSS选择器都没成功。以下是我的代码,求技术建议:

from selenium import webdriver
from selenium.webdriver.common.keys import Keys
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.service import Service

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC


options = webdriver.ChromeOptions()
options.add_argument("--log-level=OFF")
options.add_experimental_option('excludeSwitches', ['enable-logging'])

driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options)

try:
    #driver.get('https://www.worldbank.org/en/projects-operations/procurement/debarred-firms');
    #downloadcsv= driver.find_element(By.XPATH, '//*[@id="k-debarred-firms"]/div[1]/a');
    #gotit= driver.find_element(By.CLASS_NAME, "dialog_form_actions");
    #gotit.click();
    #Click on Download Button
    #driver.find_element(By.XPATH,'//*[@id="k-debarred-firms"]/div[1]/a').click()
    #time.sleep(50)

    driver.execute("get", {'url': 'https://www.worldbank.org/en/projects-operations/procurement/debarred-firms'})
    WebDriverWait(driver, 200).until(EC.element_to_be_clickable((By.CSS_SELECTOR,
'title="Excel"'))).click()

    driver.close()
    print("file downloaded")

except:
    print("Invalid URL")

问题分析与修正方案

原代码核心问题

  1. CSS选择器语法错误:title="Excel"不是合法的CSS选择器,正确写法应为a[title="Excel"],用来定位带title="Excel"属性的<a>标签
  2. 未处理页面弹窗:页面加载后会出现隐私提示弹窗,必须先关闭才能操作下载按钮
  3. 异常捕获过于宽泛:直接用except:会掩盖所有错误,无法定位具体问题
  4. 页面跳转写法冗余:driver.execute("get")不如直接用driver.get(url)简洁直观

修正后的代码

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from webdriver_manager.chrome import ChromeDriverManager
import time

# 配置Chrome选项,可指定默认下载路径(避免浏览器下载确认弹窗)
options = webdriver.ChromeOptions()
options.add_argument("--log-level=OFF")
options.add_experimental_option('excludeSwitches', ['enable-logging'])
# 替换成你的本地下载路径
prefs = {"download.default_directory": "C:/Your/Download/Path"}  
options.add_experimental_option("prefs", prefs)

driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options)

try:
    # 打开目标页面
    driver.get('https://www.worldbank.org/en/projects-operations/procurement/debarred-firms')
    
    # 等待并关闭隐私提示弹窗
    WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.CSS_SELECTOR, "button.dialog_form_actions"))
    ).click()
    
    # 等待Excel下载按钮可点击并触发点击
    excel_download_btn = WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.CSS_SELECTOR, 'a[title="Excel"]'))
    )
    excel_download_btn.click()
    
    # 等待文件下载完成(根据文件大小调整时长)
    time.sleep(10)
    
    print("文件下载完成")
except Exception as e:
    print(f"出错了: {str(e)}")
finally:
    # 确保浏览器关闭
    driver.quit()

关键说明

  • 弹窗处理:用button.dialog_form_actions定位隐私提示的确认按钮,消除后续操作的干扰
  • 精准元素定位:通过a[title="Excel"]定位下载按钮,比XPath更稳定
  • 下载路径配置:设置download.default_directory可指定文件保存位置,避免手动确认
  • 异常处理优化:捕获具体异常并打印信息,方便排查问题
  • 显式等待替代固定延时:用WebDriverWait等待元素加载,避免因页面加载慢导致的元素未找到问题

内容的提问来源于stack exchange,提问作者Raja Manikam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 17:45:28