You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于每行首个非零值重新排序Pandas DataFrame

解决方案

要实现将每行从首个非零值开始左对齐,把列改为"开业后第N个月"(或示例中的A/B/C...),可以按以下思路操作:

核心思路

  1. 定位每行第一个非零值的位置,这是门店开业的起始月份
  2. 将该行从起始位置开始的所有数值左移至DataFrame开头,前导位置留空(用NaN表示)
  3. 重新命名列,对应"开业后第1/2/...个月"

方法一:逐行处理(易理解,适合小数据集)

使用apply逐行处理,逻辑直观:

import pandas as pd
import numpy as np

# 原DataFrame
df = pd.DataFrame([[0.0, 0.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0],
                   [0.0, 0.0, 0.0, 3.0, 1.0, 1.0, 7.0, 0.0],
                   [0.0, 0.0, 0.0, 0.0, 15.0, 16.0, 17.0, 18.0]
                  ], columns=['Jan-19', 'Feb-19', 'Mar-19', 'Apr-19', 'May-19', 'June-19', 'Jul-19', 'Aug-19'])

def align_post_open(row):
    # 找到第一个非零值的列索引
    first_open_col = row.ne(0).idxmax()
    # 提取开业后的所有业绩数据
    open_data = row.loc[first_open_col:].values
    # 创建新行:前半部分补NaN,后半部分填充开业数据
    new_row = pd.Series([np.nan]*len(row))
    new_row[:len(open_data)] = open_data
    return new_row

# 应用函数到每行
result_df = df.apply(align_post_open, axis=1)
# 重命名列(A/B/C...或自定义为"开业后第1个月"等)
result_df.columns = [chr(65+i) for i in range(len(result_df.columns))]

print(result_df)

运行后输出与你期望的结果一致(NaN会显示为空):

A     B     C     D    E    F   G   H
0   3.0   4.0   5.0   6.0  7.0  8.0 NaN NaN
1   3.0   1.0   1.0   7.0  0.0  NaN NaN NaN
2  15.0  16.0  17.0  18.0  NaN  NaN NaN NaN

方法二:矢量化操作(高效,适合大数据集)

避免逐行循环,用Numpy广播实现批量处理,速度更快:

import pandas as pd
import numpy as np

df = pd.DataFrame([[0.0, 0.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0],
                   [0.0, 0.0, 0.0, 3.0, 1.0, 1.0, 7.0, 0.0],
                   [0.0, 0.0, 0.0, 0.0, 15.0, 16.0, 17.0, 18.0]
                  ], columns=['Jan-19', 'Feb-19', 'Mar-19', 'Apr-19', 'May-19', 'June-19', 'Jul-19', 'Aug-19'])

# 1. 获取每行第一个非零值的列索引(整数位置)
first_open_pos = df.ne(0).cumsum(axis=1).eq(1).idxmax(axis=1)
col_indices = df.columns.get_indexer(first_open_pos)

# 2. 构建移位矩阵,确定每个值的目标位置
row_idx = np.arange(df.shape[0])[:, None]
col_shift = np.arange(df.shape[1]) - col_indices[:, None]

# 3. 只保留开业后的有效数据,其余位置填NaN
mask = col_shift >= 0
result_values = np.where(mask, df.values[row_idx, col_shift + col_indices[:, None]], np.nan)

# 4. 生成结果DataFrame并命名列
result_df = pd.DataFrame(result_values, columns=[chr(65+i) for i in range(df.shape[1])])

print(result_df)

补充说明

  • 如果需要将NaN显示为空字符串,可以添加result_df = result_df.fillna(''),但会将数值转为字符串类型,按需选择
  • 列名可以自定义,比如改为[f"开业后第{i+1}个月" for i in range(len(result_df.columns))]

内容的提问来源于stack exchange,提问作者bbbb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 06:05:59