You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中为每个唯一Sample提取关联基因并生成Dataframe

最优实现方法

用 pandas 的 groupby 结合 agg 方法是处理这类需求的最优方案,原生高效且代码简洁:

  1. 先构造示例 DataFrame(若已有现成数据可跳过此步骤):
import pandas as pd

df = pd.DataFrame({
    'Sample': ['A', 'A', 'B', 'B', 'C', 'B', 'C'],
    'Genes': ['Gene1', 'Gene2', 'Gene3', 'Gene2', 'Gene3', 'Gene1', 'Gene4']
})
  1. 按 Sample 字段分组,将每组的 Genes 用「逗号+空格」拼接成字符串:
merged_df = df.groupby('Sample')['Genes'].agg(', '.join).reset_index()

执行后得到的 merged_df 完全符合需求格式:

SampleGenes
AGene1, Gene2
BGene3, Gene2, Gene1
CGene3, Gene4

为什么这是最优方法?

  • 属于 pandas 原生向量式操作,比手动循环遍历样本的方式效率高数倍,尤其适配大规模数据集场景
  • 代码简洁直观,核心逻辑仅一行,可读性和维护性都很强

内容的提问来源于stack exchange,提问作者user5029313

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 01:31:03