如何基于同一DataFrame的Question分组,从Answer列提取Category值生成新列?
解决方案
假设你的DataFrame名为df,可以通过以下几种方式实现需求:
方法一:分组后提取目标值并广播
利用groupby结合transform,筛选出每组中属于Category类的值(这里默认Category类值包含"Category"关键词),取第一个有效值填充整组:
import pandas as pd # 构造示例数据 data = { 'Question': [1,1,1,2,2,2], 'Answer': ['Category1','Category1','Not Important','Category2','Category2','Very Important'] } df = pd.DataFrame(data) # 新增Category列 df['Category'] = df.groupby('Question')['Answer'].transform( lambda x: x[x.str.contains('Category')].iloc[0] )
方法二:先提取分组映射关系再合并
先创建Question与对应Category的映射字典,再通过map快速填充列:
# 提取每个Question对应的Category值 category_map = df[df['Answer'].str.contains('Category')].drop_duplicates('Question').set_index('Question')['Answer'].to_dict() # 映射到原DataFrame df['Category'] = df['Question'].map(category_map)
执行后得到的结果如下:
| Question | Answer | Category |
|---|---|---|
| 1 | Category1 | Category1 |
| 1 | Category1 | Category1 |
| 1 | Not Important | Category1 |
| 2 | Category2 | Category2 |
| 2 | Category2 | Category2 |
| 2 | Very Important | Category2 |
自定义筛选规则说明
如果Category类值不是靠"Category"关键词识别,而是需要排除特定非类别值(比如"Not Important"、"Very Important"),可以修改筛选逻辑:
df['Category'] = df.groupby('Question')['Answer'].transform( lambda x: x[~x.isin(['Not Important', 'Very Important'])].iloc[0] )
内容的提问来源于stack exchange,提问作者Lisa
相关产品推荐
相关产品推荐

