如何基于时间周序列为pandas DataFrame插入缺失行并保持顺序
解决方案
你可以直接使用pandas内置的重索引/合并逻辑实现需求,不用手动循环查找缺失值和拼接,代码更简洁且天然保证周次顺序正确,同时支持同时补全多个缺失周。
短代码实现(适配你提供的同一年、固定维度的场景)
import pandas as pd # 初始化原始DF df = pd.DataFrame([[33534,9132,'Current','W41-2021',34], [33534,9132,'Current','W42-2021', 45], [33534,9132,'Current','W44-2021', 32], [33534,9132,'Current','W45-2021', 41], [33534,9132,'Current','W46-2021',49]], columns = ['Item', 'Location', 'Version', 'Time', 'Value']) # 提取周数作为临时索引并排序 df = df.set_index(df['Time'].str[1:3].astype(int)).sort_index() # 生成从最小周到最大周的完整索引 full_idx = range(df.index.min(), df.index.max()+1) # 重索引自动补全缺失行 df = df.reindex(full_idx) # 补全各列值 df['Time'] = 'W' + df.index.astype(str) + '-2021' # 固定维度列复用前后值 df[['Item','Location','Version']] = df[['Item','Location','Version']].ffill().bfill() # Value列填充0并转回整数 df['Value'] = df['Value'].fillna(0).astype(int) # 清理临时索引,恢复原有结构 df = df.reset_index(drop=True)
通用扩展版(支持跨年、多维度分组场景)
如果你的数据涉及跨年、或者Item/Location等列存在多个分组,可使用下面的通用方案:
import pandas as pd df = pd.DataFrame([[33534,9132,'Current','W41-2021',34], [33534,9132,'Current','W42-2021', 45], [33534,9132,'Current','W44-2021', 32], [33534,9132,'Current','W45-2021', 41], [33534,9132,'Current','W46-2021',49]], columns = ['Item', 'Location', 'Version', 'Time', 'Value']) # 提取周、年数值字段 df[['week_num', 'year']] = df['Time'].str.extract('W(\d+)-(\d+)').astype(int) # 按固定维度分组,每组生成完整周序列后合并 result_list = [] for (item, loc, version), group in df.groupby(['Item', 'Location', 'Version']): min_w, max_w = group['week_num'].min(), group['week_num'].max() year = group['year'].iloc[0] # 生成分组内完整周序列 full_group = pd.DataFrame({ 'week_num': range(min_w, max_w+1), 'year': year, 'Item': item, 'Location': loc, 'Version': version }) full_group['Time'] = 'W' + full_group['week_num'].astype(str) + '-' + full_group['year'].astype(str) # 合并原数据补Value full_group = full_group.merge(group[['Time', 'Value']], on='Time', how='left') result_list.append(full_group) # 合并所有分组结果,处理格式 result = pd.concat(result_list, ignore_index=True) result['Value'] = result['Value'].fillna(0).astype(int) result = result[df.columns.drop(['week_num', 'year'])]
内容的提问来源于stack exchange,提问作者PranavM
相关产品推荐
相关产品推荐

