Pandas如何删除面板数据中时间序列不完整的行
实现方案
逻辑思路
- 首先以
Country1、Country2两个字段作为分组键,区分不同的国家配对组合 - 校验每个分组的
Year字段是否完全覆盖目标时间范围(2000、2001、2002) - 保留所有校验通过的分组对应的行即可
如果你实际数据中存在同一配对同一年份有多条重复数据的情况,可以在获取年份时加
unique()去重,避免重复行影响判断结果。
代码实现
import pandas as pd # 构造示例DataFrame(你实际使用时替换为自己的df即可) data = { "Country1": ["Italy", "Italy", "Italy", "Germany", "Germany", "Mexico", "Mexico", "Mexico", "US", "US", "Greece", "Greece"], "Country2": ["Greece", "Greece", "Greece", "Italy", "Italy", "Canada", "Canada", "Canada", "France", "France", "Italy", "Italy"], "Year": [2000, 2001, 2002, 2000, 2002, 2000, 2001, 2002, 2000, 2001, 2000, 2001] } df = pd.DataFrame(data, index=range(1, 13)) # 定义需要覆盖的完整年份集合 full_years = {2000, 2001, 2002} # 分组过滤核心逻辑 filtered_df = df.groupby(["Country1", "Country2"]).filter(lambda group: set(group["Year"]) == full_years) # 可选:重置索引和你给出的示例输出序号保持一致 filtered_df = filtered_df.reset_index(drop=True) filtered_df.index += 1
输出验证
运行后filtered_df的结果和你期望的输出完全一致:
| 序号 | Country1 | Country2 | Year |
|---|---|---|---|
| 1 | Italy | Greece | 2000 |
| 2 | Italy | Greece | 2001 |
| 3 | Italy | Greece | 2002 |
| 4 | Mexico | Canada | 2000 |
| 5 | Mexico | Canada | 2001 |
| 6 | Mexico | Canada | 2002 |
内容的提问来源于stack exchange,提问作者user14237226
相关产品推荐
相关产品推荐

