You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从每日快照CSV构建历史趋势数据集?

优化CSV快照整合到历史数据集的方案

你可以直接通过文件系统检查历史文件是否存在,替代try-except的模糊判断,同时修正原代码中merge方法的逻辑问题(merge是列匹配连接,不是追加行,会导致数据丢失),以下是两种更优的实现方式:

方法一:用os模块直接判断文件存在性

import pandas as pd
import os

# 定义文件路径
snapshot_path = 'C:/source/snapshot_data.csv'
historic_path = 'C:/merged/historic_data.csv'

# 读取每日快照数据
snapshot = pd.read_csv(snapshot_path)

# 检查历史文件是否存在
if os.path.exists(historic_path):
    # 读取历史数据
    historic = pd.read_csv(historic_path)
    # 追加快照数据到历史(ignore_index重置索引避免重复)
    combined_data = pd.concat([historic, snapshot], ignore_index=True)
    # 可选:按唯一标识去重,保留最新记录(比如以'id'列为唯一键)
    # combined_data = combined_data.drop_duplicates(subset='id', keep='last')
    # 保存合并后的历史数据
    combined_data.to_csv(historic_path, index=False)
else:
    # 历史文件不存在,直接保存快照作为初始历史数据
    snapshot.to_csv(historic_path, index=False)

方法二:用pathlib(Python3.4+推荐,更现代的路径处理)

import pandas as pd
from pathlib import Path

# 用Path对象管理路径,跨平台更友好
snapshot_path = Path('C:/source/snapshot_data.csv')
historic_path = Path('C:/merged/historic_data.csv')

snapshot = pd.read_csv(snapshot_path)

if historic_path.exists():
    historic = pd.read_csv(historic_path)
    combined_data = pd.concat([historic, snapshot], ignore_index=True)
    # 按需去重
    # combined_data = combined_data.drop_duplicates(subset='unique_id', keep='last')
    combined_data.to_csv(historic_path, index=False)
else:
    snapshot.to_csv(historic_path, index=False)

关键说明

  1. 替换merge为concat:原代码的merge是做列匹配的关联操作,如果你只是想把每日快照的行追加到历史数据集中,concat才是正确的选择,避免丢失不匹配的数据。
  2. 明确逻辑分支:原代码的except未指定异常类型,会捕获所有错误(比如文件损坏、权限不足),导致误执行保存快照的操作。直接判断文件存在性,逻辑更清晰,也能区分"文件不存在"和其他读取错误。
  3. 可选去重:如果每日快照可能包含之前的历史记录,通过drop_duplicates指定唯一标识列,可以保留最新的快照数据,避免历史数据重复。

内容的提问来源于stack exchange,提问作者Joakim Torsvik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 10:05:36