如何基于Pandas DataFrame的State列条件分配对应列表值?
问题描述
现有一个包含State和Dates列的Pandas DataFrame,以及一个country列表,数据示例如下:
import pandas as pd df = pd.DataFrame({ 'State': [0,1,2,3,0,1,0,1,2,0,1,2,3,4,5], 'Dates': ['1/1/2023','2/1/2023','3/1/2023','4/1/2023', '1/1/2023','2/1/2023', '1/1/2023','2/1/2023','3/1/2023', '1/1/2023','2/1/2023','3/1/2023','4/1/2023','5/1/2023','6/1/2023'] }) country = ['A', 'B', 'C', 'D', ...]
需求:每当State列值为0时,按顺序从country列表中取出新元素,为该行及后续直至下一个State为0的行,在新增的country列中填充该元素,期望输出如下:
State Dates country 0 0 1/1/2023 A 1 1 2/1/2023 A 2 2 3/1/2023 A 3 3 4/1/2023 A 4 0 1/1/2023 B 5 1 2/1/2023 B 6 0 1/1/2023 C 7 1 2/1/2023 C 8 2 3/1/2023 C 9 0 1/1/2023 D 10 1 2/1/2023 D 11 2 3/1/2023 D 12 3 4/1/2023 D 13 4 5/1/2023 D 14 5 6/1/2023 D
解决方案
提供两种简洁高效的实现方式:
方法一:分组编号映射
通过创建分组标识,将分组与country列表元素一一对应:
# 1. 生成分组ID:每遇到State=0,分组编号递增 df['group_id'] = (df['State'] == 0).cumsum() # 2. 构建分组到国家的映射字典 group_count = df['group_id'].nunique() group_to_country = {i+1: country[i] for i in range(group_count)} # 3. 映射生成country列,删除临时分组列 df['country'] = df['group_id'].map(group_to_country) df = df.drop('group_id', axis=1)
方法二:向前填充(ffill)
直接在State=0的位置填充对应国家,再向前填充空值:
# 1. 在State=0的行填充country列表对应元素 state_zero_count = df['State'].eq(0).sum() df['country'] = None df.loc[df['State'] == 0, 'country'] = country[:state_zero_count] # 2. 向前填充空值,完成分组填充 df['country'] = df['country'].ffill()
注意事项:确保country列表的长度不小于DataFrame中State=0的行数,避免索引越界错误。
内容的提问来源于stack exchange,提问作者eeem
相关产品推荐
相关产品推荐

