You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中按列分组并依据其他列值为新列赋值的实现方法

问题描述

现有如下DataFrame:

col1        col2            col3
HP:0002616  ['HP:0001679']  Abnormal aortic morphology
HP:0002616  ['HP:0002597']  Abnormality of the vasculature
HP:0002616  ['HP:0001626']  Abnormality of the cardiovascular system
HP:0002616  ['HP:0000118']  Phenotypic abnormality
HP:0002616  ['HP:0000118']  disease
HP:0002616  ['HP:0000118']  quality
HP:0002616  ['HP:0000118']  material property
HP:0002616  ['HP:0000118']  experimental factor
HP:0002616  ['HP:0000118']  Thing
HP:0002616  ['HP:0030680']  Abnormality of cardiovascular system morphology
HP:0002616  ['HP:0002617']  Vascular dilatation 
HP:0010689  ['HP:0011297']  Abnormal digit morphology   
HP:0010689  ['HP:0011842']  Abnormal skeletal morphology    
HP:0010689  ['HP:0000924']  Abnormality of the skeletal system  
HP:0010689  ['HP:0000118']  Phenotypic abnormality  
HP:0010689  ['HP:0000118']  phenotype   

期望输出:

col1        col4           
HP:0002616  disease 
HP:0010689  phenotype   

需求规则:

  • 按col1分组
  • 若分组内col3包含'disease',col4取值'disease'
  • 若包含'phenotype',col4取值'phenotype'
  • 两者都包含时,col4取值'disease, phenotype'
解决方案

使用Pandas的groupby结合自定义函数即可实现,代码如下:

import pandas as pd

# 定义分组处理函数
def generate_col4(group):
    has_disease = 'disease' in group['col3'].values
    has_phenotype = 'phenotype' in group['col3'].values
    
    output = []
    if has_disease:
        output.append('disease')
    if has_phenotype:
        output.append('phenotype')
    
    return ', '.join(output)

# 假设原始数据存储在df变量中
result = df.groupby('col1').apply(generate_col4).reset_index(name='col4')

代码说明:

  1. 自定义函数generate_col4接收每个分组数据,检查col3中是否存在目标字符串
  2. 根据存在情况收集结果,用逗号拼接成最终的col4值
  3. 通过groupby('col1').apply()将函数应用到每个分组,最后用reset_index整理成目标DataFrame格式

内容的提问来源于stack exchange,提问作者rshar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 14:20:29