You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas多级索引中按数值逻辑排序字符串型样本名称

解决Pandas Groupby中样本名称按数值逻辑排序的问题

核心思路

问题本质是Pandas对字符串类型的分组键默认采用字典序排序,要实现自然数值排序,关键是让Pandas识别样本名称中的数值顺序——通过有序分类(Categorical)或自定义排序索引来控制分组结果的顺序,同时完整保留分组结构。

方法一:提取数值生成有序分类(推荐)

先从样本名称中提取数字部分,用它生成正确排序的分类列,再执行分组计算:

import pandas as pd

# 1. 从Sample Name中提取数字并转为整数(适配多位数如10、100)
df['Sample_Num'] = df['Sample Name'].str.extract(r'(\d+)').astype(int)

# 2. 按Assay Num和提取的数值排序,得到每个检测编号下的正确样本顺序
sorted_sample_order = df.sort_values(['Assay Num', 'Sample_Num'])['Sample Name'].unique()

# 3. 将Sample Name转为有序分类,强制指定排序规则
df['Sample Name'] = pd.Categorical(df['Sample Name'], categories=sorted_sample_order, ordered=True)

# 4. 执行分组计算,结果会按数值逻辑排序
grouped_result = df.groupby(['Assay Num', 'Sample Name'], sort=True)['Resp 1'].mean().reset_index()

说明

  • 正则(\d+)能精准捕获样本名称中的连续数字,避免单/多位数排序混乱
  • 有序分类会完全替代Pandas默认的字典序,确保分组结果按我们定义的顺序输出

方法二:用natsort库直接生成自然排序索引

如果不想额外生成数值列,可借助natsort库直接对分组后的索引做自然排序:

from natsort import index_natsorted
import pandas as pd

# 1. 先正常执行分组计算
grouped_result = df.groupby(['Assay Num', 'Sample Name'])['Resp 1'].mean()

# 2. 对多级索引的Sample Name部分应用自然排序,生成正确的索引顺序
sorted_index = grouped_result.index.sortlevel(level='Assay Num')[0].sort_values(
    level='Sample Name',
    key=lambda x: index_natsorted(x)
)

# 3. 重新索引结果,得到符合数值逻辑的排序
grouped_result = grouped_result.reindex(sorted_index).reset_index()

说明

  • index_natsorted返回自然排序的索引位置,适合作为sort_values的key参数
  • 先按Assay Num排序,再对每个检测编号下的样本做自然排序,保证分组结构不被破坏

方法三:自定义多级分组键分类

如果需要严格保证每个检测编号下的样本独立排序,可以直接生成包含检测编号和样本名称的有序分类:

from natsort import natsorted
from pandas.api.types import CategoricalDtype
import pandas as pd

# 1. 遍历每个检测编号,对其下的样本名称做自然排序
sorted_group_keys = []
for assay_num in df['Assay Num'].unique():
    sample_names = df[df['Assay Num'] == assay_num]['Sample Name'].unique()
    sorted_samples = natsorted(sample_names)
    sorted_group_keys.extend([(assay_num, name) for name in sorted_samples])

# 2. 生成多级分类类型
group_cat_type = CategoricalDtype(categories=sorted_group_keys, ordered=True)

# 3. 将分组键转为元组并设置为分类
df['Group_Key'] = list(zip(df['Assay Num'], df['Sample Name']))
df['Group_Key'] = df['Group_Key'].astype(group_cat_type)

# 4. 分组计算后还原原始列
grouped_result = df.groupby('Group_Key')['Resp 1'].mean().reset_index()
grouped_result[['Assay Num', 'Sample Name']] = pd.DataFrame(
    grouped_result['Group_Key'].tolist(),
    index=grouped_result.index
)
grouped_result = grouped_result.drop('Group_Key', axis=1)

说明

  • 这种方式完全自定义每个(检测编号, 样本名称)组合的顺序,适合复杂的分组排序需求
  • 最后通过拆分元组还原原始分组列,不影响最终结果的结构

内容的提问来源于stack exchange,提问作者MixMaster32395

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 16:35:24