You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python按月份重采样筛选与target列MAPE最小的两列并处理数据

实现代码
import pandas as pd
import numpy as np

# 1. 生成测试数据
np.random.seed(2021)
dates = pd.date_range('20130226', periods=90)
df = pd.DataFrame(np.random.uniform(0, 10, size=(90, 4)), index=dates, columns=['A_values', 'B_values', 'C_values', 'target'])

# 2. 定义MAPE计算函数
def mape(y_true, y_pred):
    y_true, y_pred = np.array(y_true), np.array(y_pred)
    return np.mean(np.abs((y_true - y_pred) / np.maximum(np.ones(len(y_true)), np.abs(y_true))))*100

# 3. 定义单月MAPE计算逻辑:输入单月分组数据,返回三个特征列的MAPE值
def calc_month_mape(month_df):
    return pd.Series({
        'A_values': mape(month_df['target'], month_df['A_values']),
        'B_values': mape(month_df['target'], month_df['B_values']),
        'C_values': mape(month_df['target'], month_df['C_values'])
    })

# 4. 按月份重采样,计算每个月三个特征列的MAPE
monthly_mape = df.resample('M').apply(calc_month_mape)

# 5. 筛选每个月MAPE最小的2个特征列
month_keep_cols = monthly_mape.apply(lambda x: x.nsmallest(2).index.tolist(), axis=1)

# 6. 对原始数据按月份规则置空非最小误差列
feature_cols = ['A_values', 'B_values', 'C_values']
# 生成每个日期对应月份的月末标识,用于匹配上面的月度规则
df['month_key'] = df.index + pd.offsets.MonthEnd(0)
for col in feature_cols:
    # 标记当前列在对应月份是否需要保留
    keep_mask = df['month_key'].map(lambda k: col in month_keep_cols.loc[k])
    df.loc[~keep_mask, col] = np.nan
# 删除辅助计算的月份标识列
df = df.drop('month_key', axis=1)

# 查看结果
print(df.head(10))

注意说明

  • 原有伪代码的错误原因是直接传入全量列计算MAPE,没有实现按分组计算的逻辑,需要将分组内的计算逻辑封装为函数传入apply才能得到每个月的独立MAPE结果。
  • 代码里使用nsmallest(2)直接取每个月MAPE最小的两个列名,不需要手动对比三个列的数值。
  • 最后通过月份匹配的方式给原始数据打掩码,仅保留指定列的原始值,其余列置为NaN。

内容的提问来源于stack exchange,提问作者ah bon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 20:45:08