如何用Python自动每5分钟读取CSV文件并应用Pandas自定义函数?
自动化CSV文件的定时读取与Pandas处理流程
嘿,作为编程初学者,想要摆脱手动重复操作的烦恼完全没问题!我来帮你把这段手动流程改成自动化的,让它每5分钟自动运行,不用再手动命名文件和重复写代码~
先看看你当前的代码问题:每次都要手动指定文件名、创建新变量,重复调用预处理和合并函数,这不仅繁琐还容易出错。下面是一步步的解决方案:
1. 自动识别目标CSV文件
不用硬编码每个文件名,我们可以用glob模块匹配所有符合格式的CSV文件,还能按日期排序确保处理顺序正确:
import glob import pandas as pd # 匹配所有类似foDDMAYYYYbhav.csv格式的文件,并按文件名排序(保证日期顺序) csv_files = sorted(glob.glob('fo*MAY2018bhav.csv'))
2. 批量处理并逐步合并DataFrame
把重复的预处理、合并操作改成循环,让代码自动处理所有文件,不用手动创建df_9May、df_10May这类变量:
# 初始化合并后的结果DataFrame combined_df = None for file in csv_files: # 读取CSV并解析日期 df = pd.read_csv(file, parse_dates=True) # 应用你的自定义预处理函数 df = PreprocessDataframe(df) if combined_df is None: # 第一个文件直接作为初始结果 combined_df = df else: # 合并当前文件与之前的结果 combined_df = combineDFs(combined_df, df) # 应用你的NetVal和均价计算函数 combined_df = NewNetVal_AvgPrice(combined_df)
这样不管有多少个CSV文件,都会按顺序自动处理、合并,完全不用手动干预。
3. 设置每5分钟自动运行的定时任务
要实现定时执行,推荐用schedule库(比直接用time.sleep更灵活易读)。先安装它:
pip install schedule
然后编写定时任务代码:
import schedule import time def run_automated_process(): # 把上面的文件识别和处理逻辑放在这里 csv_files = sorted(glob.glob('fo*MAY2018bhav.csv')) combined_df = None for file in csv_files: df = pd.read_csv(file, parse_dates=True) df = PreprocessDataframe(df) if combined_df is None: combined_df = df else: combined_df = combineDFs(combined_df, df) combined_df = NewNetVal_AvgPrice(combined_df) # 处理完成后可以保存结果或输出日志 print(f"Process completed at {pd.Timestamp.now()}") combined_df.to_csv('final_combined_data.csv', index=False) # 设置每5分钟运行一次任务 schedule.every(5).minutes.do(run_automated_process) # 保持脚本持续运行,等待定时任务触发 while True: schedule.run_pending() time.sleep(1)
额外优化:避免重复处理已读文件
上面的代码每次运行都会重新处理所有文件,如果你希望只处理新增的CSV,可以记录已处理的文件名,跳过重复文件:
# 用集合记录已处理的文件名,避免重复操作 processed_files = set() def run_automated_process(): global processed_files csv_files = sorted(glob.glob('fo*MAY2018bhav.csv')) combined_df = None for file in csv_files: if file in processed_files: continue # 跳过已经处理过的文件 df = pd.read_csv(file, parse_dates=True) df = PreprocessDataframe(df) if combined_df is None: combined_df = df else: combined_df = combineDFs(combined_df, df) combined_df = NewNetVal_AvgPrice(combined_df) processed_files.add(file) # 标记为已处理 if combined_df is not None: combined_df.to_csv('final_combined_data.csv', index=False) print(f"New data processed at {pd.Timestamp.now()}")
这样你的自动化流程就完全搭建好了,运行脚本后它会每5分钟自动检查并处理新的CSV文件,再也不用手动重复操作啦!
内容的提问来源于stack exchange,提问作者Mohit Aneja
相关产品推荐
相关产品推荐

