Python Pandas多索引reindex:按id补全日期时country列全NaN问题
问题根因
代码返回country列全NaN由两个核心问题导致:
- 类型不匹配:原始DataFrame的
date列是字符串格式,pd.date_range生成的是datetime64时间类型,reindex时索引值类型不一致,完全匹配不到原始行数据,所有值初始都是空。 - 填充逻辑错误:即使对齐类型,直接全局
ffill会跨id填充值,比如A组的国家值会错误填充到B组的行,必须按id分组做前向填充。
修复代码
先把date列转为datetime类型,重排索引后分组填充即可:
import pandas as pd test = {"id": ["A", "A", "A", "B", "B", "B"], "date": ["09-02-2013", "09-03-2013", "09-05-2013", "09-15-2013", "09-17-2013", "09-18-2013"], "country": ["Poland", "Poland", "France", "Scotland", "Scotland", "Canada"]} test_df = pd.DataFrame(test) # 转换日期列为datetime类型,统一索引格式 test_df["date"] = pd.to_datetime(test_df["date"], format="%m-%d-%Y") # 计算每个id对应的日期上下限 date_bounds = test_df.groupby("id")["date"].agg(["min", "max"]) # 构造补全后的(date, id)多重索引 full_idx = pd.MultiIndex.from_frame( date_bounds.apply(lambda x: pd.date_range(x["min"], x["max"], freq="D"), axis=1) .explode() .reset_index(name="date")[["date", "id"]] ) # 重排数据 test_df = test_df.set_index(["date", "id"]).reindex(full_idx) # 按id分组前向填充,避免跨id填错值 test_df["country"] = test_df.groupby(level="id")["country"].ffill() # 重置索引,如需保留原日期字符串格式可加这步 test_df = test_df.reset_index() test_df["date"] = test_df["date"].dt.strftime("%m-%d-%Y")
运行后输出结果完全符合预期:
| id | date | country |
|---|---|---|
| A | 09-02-2013 | Poland |
| A | 09-03-2013 | Poland |
| A | 09-04-2013 | Poland |
| A | 09-05-2013 | France |
| B | 09-15-2013 | Scotland |
| B | 09-16-2013 | Scotland |
| B | 09-17-2013 | Scotland |
| B | 09-18-2013 | Canada |
更简洁的实现
pandas 1.3及以上版本可以直接用分组重采样实现,不需要手动构造索引,代码更简洁不易出错:
test_df["date"] = pd.to_datetime(test_df["date"], format="%m-%d-%Y") result = ( test_df.set_index("date") .groupby("id") .resample("D") .ffill() .reset_index() ) # 转回原字符串日期格式 result["date"] = result["date"].dt.strftime("%m-%d-%Y")
内容的提问来源于stack exchange,提问作者FilipaF
相关产品推荐
相关产品推荐

