You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python按年月分组删除全NaN列并重命名为top_k格式的实现方法

实现方案

你提到的justify_nd是scipy提供的数组对齐工具,这个需求用pandas原生功能就能实现,逻辑更清晰无额外依赖,以下是可直接运行的完整方案:


1. 依赖导入&测试数据构造

先构造和你场景匹配的测试数据,方便直接验证效果:

import pandas as pd
import numpy as np

# 测试数据结构:日期索引、带_mape后缀的特征列、target列
date_rng = pd.date_range(start='2023-01-01', end='2023-02-28', freq='D')
df = pd.DataFrame(
    {
        'feat1_mape': [np.nan]*31 + [2.1,3.2,1.5]*9 + [np.nan], # 1月全空、2月有值
        'feat2_mape': [1.2,2.3,0.8]*10 + [np.nan] + [np.nan]*28, # 1月有值、2月全空
        'feat3_mape': [0.5,1.1,0.7]*10 + [np.nan] + [3.3,2.2,4.1]*9 + [np.nan], # 两个月都有值
        'target': np.random.randint(10, 100, size=len(date_rng))
    },
    index=date_rng
)

2. 核心处理逻辑

# 拆分MAPE特征列和target列
mape_cols = [col for col in df.columns if col.endswith('_mape')]
target_s = df['target']
mape_df = df[mape_cols]

# 定义单月分组处理函数
def process_month_group(group):
    # 删除当前组内全为NaN的列
    valid_group = group.dropna(axis=1, how='all')
    # 把剩余有效列重命名为top_1、top_2...格式
    valid_group.columns = [f'top_{i+1}' for i in range(len(valid_group.columns))]
    return valid_group

# 按年月分组应用处理逻辑
processed_mape = mape_df.groupby(mape_df.index.to_period('M'), group_keys=False).apply(process_month_group)

# 合并处理后的特征和原target列得到最终结果
final_df = pd.concat([processed_mape, target_s], axis=1)

可选:统一所有行的列数

如果需要所有月份的列数对齐,不足的位置补NaN,可以在合并前加以下逻辑:

# 计算最大有效列数,生成统一列名
max_col_count = processed_mape.columns.str.extract(r'top_(\d+)').astype(int).max()[0]
unified_cols = [f'top_{i+1}' for i in range(max_col_count)]
# 重对齐列,缺失列自动补NaN
processed_mape = processed_mape.reindex(columns=unified_cols)

内容的提问来源于stack exchange,提问作者ah bon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 22:36:02