You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas提取Primary Mod Site列重复项并保留对应列最大值?

解决方案

原代码存在的问题

  • groupby("gene") 参数错误,应该用列名 "Primary Mod Site" 而非变量 gene
  • 仅拼接了重复组,但未实现保留对应化合物列最大值的核心需求

正确实现代码

假设Excel中B-M列对应数值型的化合物数据列,我们可以通过两种方式实现需求:

方式1:分组聚合取最大值(优先推荐)

import pandas as pd

# 读取Excel文件
df = pd.read_excel("20220825_CISLIB01_Plate-1_Rows-A-B")

# 定义分组键和需要取最大值的列(B-M列,按列位置索引)
group_key = "Primary Mod Site"
value_cols = df.columns[1:13]  # B列对应索引1,M列对应索引12,左闭右开取1到13

# 分组后对目标列取最大值,保留分组键
result = df.groupby(group_key)[value_cols].max().reset_index()

# 如果需要保留其他非数值列(比如Compound Name),可以结合first()聚合
# result = df.groupby(group_key).agg({
#     **{col: 'max' for col in value_cols},
#     'Compound Name': 'first'  # 保留每组第一个化合物名称,可按需调整
# }).reset_index()

# 输出或保存结果
print(result)
# result.to_excel("processed_data.xlsx", index=False)

方式2:排序后去重

如果需要保留原始数据的完整行结构(即对应最大值的整行数据),可以先排序再去重:

import pandas as pd

df = pd.read_excel("20220825_CISLIB01_Plate-1_Rows-A-B")

group_key = "Primary Mod Site"
value_cols = df.columns[1:13]

# 按分组键升序、化合物列降序排序,确保每组最大值行排在最前
df_sorted = df.sort_values(
    by=[group_key] + list(value_cols),
    ascending=[True] + [False]*len(value_cols)
)

# 保留每组第一行(即最大值行)
result = df_sorted.drop_duplicates(subset=group_key, keep="first").reset_index(drop=True)

print(result)

关键说明

  • df.columns[1:13] 是按列位置选取B-M列,如果Excel列名有明确标识,也可以直接写列名列表(比如 ["Comp_01", "Comp_02", ..., "Comp_12"])
  • 如果只需针对某一列取最大值,直接把 value_cols 指定为该列名即可

内容的提问来源于stack exchange,提问作者dominique

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 03:25:12