使用Selenium批量下载URL文件并重命名:循环匹配问题求助
解决WSJ股票历史数据下载后逐次重命名的问题
问题描述
尝试通过tickers列表生成URL下载股票历史价格数据,但当前代码会先遍历完所有URL完成全部下载,再统一重命名,导致报错(因为下载的文件都是同名HistoricalPrices.csv,后续下载会覆盖之前的,重命名时找不到对应文件)。需要修改代码,实现每次循环处理一个ticker,下载完成后立即重命名,再进入下一轮循环。
修改后的完整代码
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By import os import time import pandas as pd import datetime from datetime import datetime start = '10/26/2020' end = '1/22/2023' tickers = ["ap","dmc","creit","chib","fli","fb","dmc","fph","gma7","ltg", "mbt","mreit","nikl","pse","rcr","rlc","rrhi","scc","secb"] # 修正原tickers列表中",mreit"的多余逗号问题 path = "/Users/sef/Documents/Py-MSC/chromedriver_mac_arm64/chromedriver" folder = "/Users/sef/Documents/PSE_Data Repository" options = Options() options.add_experimental_option('detatch', True) chromeOptions = webdriver.ChromeOptions() prefs = {"download.default_directory" : folder} chromeOptions.add_experimental_option("prefs", prefs) driver = webdriver.Chrome(service=Service(path), options=chromeOptions) wait = WebDriverWait(driver, 10) # 直接遍历tickers列表,实现ticker与下载操作一一对应 for ticker in tickers: url = f'https://www.wsj.com/market-data/quotes/PH/{ticker}/historical-prices' driver.get(url) # 等待日期输入框加载完成,替代固定sleep提升稳定性 beg_date = wait.until(EC.presence_of_element_located((By.ID, "selectDateFrom"))) beg_date.clear() beg_date.send_keys(start) end_date = wait.until(EC.presence_of_element_located((By.ID, "selectDateTo"))) end_date.clear() end_date.send_keys(end) # 点击日期确认按钮 wait.until(EC.element_to_be_clickable((By.ID, "datPickerButton"))).click() # 等待下载按钮可点击并触发下载 wait.until(EC.element_to_be_clickable((By.ID, "dl_spreadsheet"))).click() # 等待下载文件生成,避免未完成就执行重命名 download_file = os.path.join(folder, "HistoricalPrices.csv") while not os.path.exists(download_file): time.sleep(1) # 立即重命名当前下载的文件 label = ticker.upper() new_file = os.path.join(folder, f"{label}.csv") # 若目标文件已存在则先删除,避免重命名报错 if os.path.exists(new_file): os.remove(new_file) os.rename(download_file, new_file) driver.quit()
关键修改说明
- 合并循环逻辑:不再单独生成
urls列表,直接遍历tickers,确保每次循环只处理一个ticker,下载与重命名操作绑定。 - 重命名移至循环内部:下载完成后立即执行重命名,彻底避免后续下载覆盖当前的
HistoricalPrices.csv文件。 - 优化等待机制:用
WebDriverWait替代固定time.sleep(),等待元素加载完成再操作,减少无效等待并提升代码稳定性。 - 增加文件校验:重命名前检查目标文件是否存在,避免因文件已存在报错;同时等待下载文件生成后再执行重命名,防止操作未完成的文件。
内容的提问来源于stack exchange,提问作者Gabriel Yousef Ramos
相关产品推荐
相关产品推荐

