You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于前后值填充pandas DataFrame中间缺失的连续数值?

问题描述

现有如下pandas DataFrame:

col     col2
0    0  Repeat1
1    3  Repeat2
2    5  Repeat3
3    7  Repeat4
4    9  Repeat5

可通过以下代码复现:

L= [0,3,5,7,9]
L2 = ['Repeat1','Repeat2','Repeat3','Repeat4','Repeat5']

import pandas as pd
df = pd.DataFrame({'col':L})
df['col2']= L2
print(df)

需要填充col列中缺失的中间连续数值,得到如下结果:

col     col2
0    0  Repeat1
1    1  Repeat1
2    2  Repeat1
3    3  Repeat2
4    4  Repeat2
5    5  Repeat3
6    6  Repeat3
7    7  Repeat4
8    8  Repeat4
9    9  Repeat5

希望找到更简洁的函数式实现方法。

解决方案

方法一:辅助列+explode(可读性优先)

利用shift确定每个分组的数值区间终点,结合apply生成连续序列后展开,逻辑清晰且代码简洁:

import pandas as pd

# 生成每个Repeat对应的数值区间终点
df['end'] = df['col'].shift(-1).fillna(df['col'].iloc[-1]).astype(int)
# 生成连续序列并展开,清理辅助列
result = df.assign(col=df.apply(lambda x: range(x['col'], x['end'] + 1), axis=1)) \
           .explode('col') \
           .drop('end', axis=1) \
           .reset_index(drop=True)

print(result)

方法二:无辅助列精简版(代码更紧凑)

无需额外辅助列,直接在apply中通过索引定位下一个数值,实现一行式函数式处理:

result = df.assign(
    col=df.apply(
        lambda x: range(x['col'], df['col'].iloc[x.name+1] if x.name < len(df)-1 else x['col'] + 1),
        axis=1
    )
).explode('col').reset_index(drop=True)

核心逻辑说明

两种方法均基于函数式思想:

  1. 对每一行生成对应col2分组的连续整数序列;
  2. 用explode将序列拆分为独立行,自动继承对应col2的值;
  3. 最终重置索引得到目标结构。

内容的提问来源于stack exchange,提问作者Bhargav

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 18:10:31