You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas按line_id分组并生成annotation_list列?

问题场景

给定名为annotations的DataFrame:

line_id  start  end  label
annotation_id                              
27229_36603_2_3        6    321  340    LOC
4                      8    200  203    PER
6                      8    262  268    PER
7                      8    262  281    ORG
10                    10     35   38    PER
12                    11    146  155    ORG
13                    11    156  164    ORG

需要按line_id分组,每个line_id对应一行,新增annotation_list列,该列是由每行start、end、label组成子列表的嵌套列表,期望输出:

line_id   annotation_list

      6   [[321, 340, 'LOC']]
      8   [[200, 203, 'PER'], [262, 268, 'PER'], [262, 281, 'ORG']]
     10   [[35, 38, 'PER']]
     11   [[146, 155, 'ORG'], [156, 164, 'ORG']]
解决方案

通过groupby结合apply实现自定义多列聚合,核心代码如下:

import pandas as pd

# 若已有annotations可跳过构造步骤
data = {
    'line_id': [6, 8, 8, 8, 10, 11, 11],
    'start': [321, 200, 262, 262, 35, 146, 156],
    'end': [340, 203, 268, 281, 38, 155, 164],
    'label': ['LOC', 'PER', 'PER', 'ORG', 'PER', 'ORG', 'ORG']
}
annotations = pd.DataFrame(data, index=['27229_36603_2_3', '4', '6', '7', '10', '12', '13'])
annotations.index.name = 'annotation_id'

# 分组生成嵌套列表
result = annotations.groupby('line_id').apply(
    lambda x: x[['start', 'end', 'label']].values.tolist()
).reset_index(name='annotation_list')

print(result)
代码说明
  • groupby('line_id'):按line_id对原DataFrame完成分组
  • apply(lambda x: ...):针对每个分组的子DataFrame,提取start/end/label三列,通过values.tolist()将每一行转为子列表,最终汇总成嵌套列表
  • reset_index(name='annotation_list'):将分组的line_id从索引转为普通列,并为新生成的嵌套列表列命名为annotation_list

此前用aggregate未成功,是因为agg默认针对单列做聚合,无法直接处理多列打包成嵌套列表的自定义逻辑,apply更适配这类需求。


内容的提问来源于stack exchange,提问作者Daniel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 14:10:06