You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python结合Selenium和Requests自动化批量下载CSV文件

解决CSV自动化批量下载问题

你遇到的核心问题是requests请求未携带登录后的会话Cookie,网站权限验证未通过,因此返回HTML而非CSV文件。以下是具体修正方案:

关键思路

Selenium完成登录后,浏览器已保存登录状态的Cookie,将这些Cookie同步到requests的会话中,就能让requests请求带上合法登录权限,成功获取CSV内容。

修正后的代码

from selenium import webdriver
import requests

# Selenium登录并导航获取下载URL
driver = webdriver.Chrome()
driver.maximize_window()
driver.get(login_url)
driver.find_element('name','email').send_keys(u_name)
driver.find_element('name','pass').send_keys(pw)
driver.find_element('xpath','/html/body/div[1]/div/div/form/dl/dd[3]/button').click()

ele = driver.find_element('xpath',
                          "/html/body/div[1]/article/div[2]/section[4]/ul/li[1]/a")
cl_url = ele.get_attribute('href')
driver.get(cl_url)

xpath_idx = '/html/body/div[1]/div/ul/li['
xpath_suffix = ']/a'
xpath_ovr = xpath_idx + '10' + xpath_suffix
driver.find_element('xpath',xpath_ovr).click()

driver.find_element('xpath', '/html/body/div[1]/div/div[10]/span[1]/a').click()
element = driver.find_element('xpath',
                              '/html/body/div[1]/div/div[2]/section[4]/div/div[2]/div/table/tbody/tr/td[1]/div/div[2]/div/a[1]')
final_url = element.get_attribute('href')

# 同步Selenium的Cookie到requests会话
session = requests.Session()
for cookie in driver.get_cookies():
    session.cookies.set(cookie['name'], cookie['value'])

# 请求并保存CSV文件
response = session.get(final_url)
# 验证响应是否为CSV类型
if 'text/csv' in response.headers.get('Content-Type', ''):
    # 从URL提取id作为文件名,避免重名
    file_id = final_url.split('id=')[1].split('&')[0]
    file_name = f"mutation_{file_id}.csv"
    with open(file_name, 'wb') as f:
        f.write(response.content)
    print(f"文件 {file_name} 已成功保存")
else:
    print("请求未返回CSV内容,请检查URL或登录状态")

driver.quit()

批量下载扩展方案

如果要批量处理多个下载URL:

  1. 将所有需要下载的URL收集到列表,如url_list = [url_1, url_2, ..., url_n]
  2. 在获取登录会话后,循环遍历列表,重复执行请求和保存步骤
  3. 添加适当延时(如import time; time.sleep(1)),避免请求过于频繁触发网站反爬机制

额外优化建议

  • 直接使用完整的final_url请求,不要拆分参数拼接,避免出现编码或参数丢失问题
  • 保存时使用response.content而非response.text,防止编码错误破坏CSV格式
  • 可通过响应头自动提取官方文件名:
    from urllib.parse import unquote
    content_disposition = response.headers.get('Content-Disposition', '')
    if 'filename=' in content_disposition:
        file_name = unquote(content_disposition.split('filename=')[1])
    

内容的提问来源于stack exchange,提问作者Siva kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 06:30:45