如何创建create_col列?基于水果历史记录判断wanted字段状态
实现创建
create_col列的方案 需求说明
针对每行数据,找到对应水果上次出现的记录,若该记录的wanted值为yes,则create_col填True;否则留空。
示例数据
| wanted | fruit | create_col | 说明 |
|---|---|---|---|
| yes | apple | 无上次记录,留空 | |
| pear | 无上次记录,留空 | ||
| pear | 上次pear的wanted不为yes,留空 | ||
| apple | True | 上次apple的wanted为yes,填True |
代码实现(基于Pandas)
import pandas as pd # 构造示例数据 data = { 'wanted': ['yes', '', '', ''], 'fruit': ['apple', 'pear', 'pear', 'apple'] } df = pd.DataFrame(data) # 按fruit分组,获取每组的上一行wanted值 df['prev_wanted'] = df.groupby('fruit')['wanted'].shift(1) # 根据规则填充create_col:prev_wanted为'yes'则填True,否则为缺失值 df['create_col'] = df['prev_wanted'].apply(lambda x: True if x == 'yes' else pd.NA) # 可选:删除中间辅助列 df = df.drop('prev_wanted', axis=1) print(df)
输出结果
wanted fruit create_col 0 yes apple <NA> 1 pear <NA> 2 pear <NA> 3 apple True
逻辑说明
- 分组移位:通过
groupby('fruit')['wanted'].shift(1)对每个水果的wanted列向上移位,得到每行对应水果的上一次wanted值。 - 条件填充:对移位后的结果做判断,仅当值为
yes时填充True,其余情况设为Pandas标准缺失值pd.NA,对应示例中的空值状态。
内容的提问来源于stack exchange,提问作者arv
相关产品推荐
相关产品推荐

