You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于同一DataFrame的Question分组,从Answer列提取Category值生成新列?

解决方案

假设你的DataFrame名为df,可以通过以下几种方式实现需求:

方法一:分组后提取目标值并广播

利用groupby结合transform,筛选出每组中属于Category类的值(这里默认Category类值包含"Category"关键词),取第一个有效值填充整组:

import pandas as pd

# 构造示例数据
data = {
    'Question': [1,1,1,2,2,2],
    'Answer': ['Category1','Category1','Not Important','Category2','Category2','Very Important']
}
df = pd.DataFrame(data)

# 新增Category列
df['Category'] = df.groupby('Question')['Answer'].transform(
    lambda x: x[x.str.contains('Category')].iloc[0]
)

方法二:先提取分组映射关系再合并

先创建Question与对应Category的映射字典,再通过map快速填充列:

# 提取每个Question对应的Category值
category_map = df[df['Answer'].str.contains('Category')].drop_duplicates('Question').set_index('Question')['Answer'].to_dict()

# 映射到原DataFrame
df['Category'] = df['Question'].map(category_map)

执行后得到的结果如下:

QuestionAnswerCategory
1Category1Category1
1Category1Category1
1Not ImportantCategory1
2Category2Category2
2Category2Category2
2Very ImportantCategory2

自定义筛选规则说明

如果Category类值不是靠"Category"关键词识别,而是需要排除特定非类别值(比如"Not Important"、"Very Important"),可以修改筛选逻辑:

df['Category'] = df.groupby('Question')['Answer'].transform(
    lambda x: x[~x.isin(['Not Important', 'Very Important'])].iloc[0]
)

内容的提问来源于stack exchange,提问作者Lisa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 10:22:08