You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何编写函数实现DataFrame时间范围筛选并生成清洗后数据集?

解决方法

你的代码有两个关键问题:

  • 函数参数重复定义了name,这会直接触发语法错误
  • 试图通过'df'_ + name这种方式动态创建变量名,Python不支持这种赋值语法,而且这种写法会让代码难以维护

正确的做法是让函数直接返回处理后的DataFrame,然后你自己把返回值赋值给想要的变量名即可,不需要在函数里搞动态变量名。

正确的函数实现

import pandas as pd

def filter_by_time_range(df, months_offset, date_col='Date'):
    # 计算时间范围起始点:最大日期往前推指定月份数
    start_time = df[date_col].max() - pd.DateOffset(months=months_offset)
    # 返回筛选后的新DataFrame
    return df.loc[df[date_col] >= start_time]

使用示例

对应你原来的df_a处理逻辑,现在可以这么用:

# 筛选df_a中最近12个月的数据(max日期减11个月,覆盖最近12个月范围)
df_a_12m = filter_by_time_range(df_a, months_offset=11)

# 处理另一个DataFrame df_b,筛选最近6个月的数据
df_b_6m = filter_by_time_range(df_b, months_offset=5)

# 如果日期列不是默认的'Date',可以指定列名
df_c_3m = filter_by_time_range(df_c, months_offset=2, date_col='交易日期')

这样既避免了重复代码,又能灵活处理不同的DataFrame和日期列,变量名由你自己控制,清晰易读。

内容的提问来源于stack exchange,提问作者Steve

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 10:48:16