You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium批量下载URL文件并重命名:循环匹配问题求助

解决WSJ股票历史数据下载后逐次重命名的问题

问题描述

尝试通过tickers列表生成URL下载股票历史价格数据,但当前代码会先遍历完所有URL完成全部下载,再统一重命名,导致报错(因为下载的文件都是同名HistoricalPrices.csv,后续下载会覆盖之前的,重命名时找不到对应文件)。需要修改代码,实现每次循环处理一个ticker,下载完成后立即重命名,再进入下一轮循环。

修改后的完整代码

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
import os
import time
import pandas as pd
import datetime
from datetime import datetime

start = '10/26/2020'
end = '1/22/2023'
tickers = ["ap","dmc","creit","chib","fli","fb","dmc","fph","gma7","ltg",
           "mbt","mreit","nikl","pse","rcr","rlc","rrhi","scc","secb"]

# 修正原tickers列表中",mreit"的多余逗号问题
path = "/Users/sef/Documents/Py-MSC/chromedriver_mac_arm64/chromedriver"
folder = "/Users/sef/Documents/PSE_Data Repository"

options = Options()
options.add_experimental_option('detatch', True)
chromeOptions = webdriver.ChromeOptions()
prefs = {"download.default_directory" : folder}
chromeOptions.add_experimental_option("prefs", prefs)

driver = webdriver.Chrome(service=Service(path), options=chromeOptions)
wait = WebDriverWait(driver, 10)

# 直接遍历tickers列表,实现ticker与下载操作一一对应
for ticker in tickers:
    url = f'https://www.wsj.com/market-data/quotes/PH/{ticker}/historical-prices'
    driver.get(url)
    
    # 等待日期输入框加载完成,替代固定sleep提升稳定性
    beg_date = wait.until(EC.presence_of_element_located((By.ID, "selectDateFrom")))
    beg_date.clear()
    beg_date.send_keys(start)
    
    end_date = wait.until(EC.presence_of_element_located((By.ID, "selectDateTo")))
    end_date.clear()
    end_date.send_keys(end)
    
    # 点击日期确认按钮
    wait.until(EC.element_to_be_clickable((By.ID, "datPickerButton"))).click()
    # 等待下载按钮可点击并触发下载
    wait.until(EC.element_to_be_clickable((By.ID, "dl_spreadsheet"))).click()
    
    # 等待下载文件生成,避免未完成就执行重命名
    download_file = os.path.join(folder, "HistoricalPrices.csv")
    while not os.path.exists(download_file):
        time.sleep(1)
    
    # 立即重命名当前下载的文件
    label = ticker.upper()
    new_file = os.path.join(folder, f"{label}.csv")
    # 若目标文件已存在则先删除,避免重命名报错
    if os.path.exists(new_file):
        os.remove(new_file)
    os.rename(download_file, new_file)

driver.quit()

关键修改说明

  • 合并循环逻辑:不再单独生成urls列表,直接遍历tickers,确保每次循环只处理一个ticker,下载与重命名操作绑定。
  • 重命名移至循环内部:下载完成后立即执行重命名,彻底避免后续下载覆盖当前的HistoricalPrices.csv文件。
  • 优化等待机制:用WebDriverWait替代固定time.sleep(),等待元素加载完成再操作,减少无效等待并提升代码稳定性。
  • 增加文件校验:重命名前检查目标文件是否存在,避免因文件已存在报错;同时等待下载文件生成后再执行重命名,防止操作未完成的文件。

内容的提问来源于stack exchange,提问作者Gabriel Yousef Ramos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 20:30:55