You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

按label分组统计content含多字段的唯一api_spec_id数量

解决方案

步骤说明

  1. 标记每条数据的content字段是否包含多个逗号分隔项(即至少存在1个逗号);
  2. 按label分组,筛选出满足条件的行,统计每组内唯一api_spec_id的数量;
  3. 补全所有需要统计的label类别(包括Nan、minor、patch、major),缺失类别计数设为0。

代码实现

import pandas as pd

# 构建示例DataFrame(若已有现成DataFrame可跳过此步)
data = {
    'label': [None, None, None, None, None, 'minor', 'patch', 'patch', 'minor', 'minor', 'major'],
    'api_spec_id': [375.0, 375.0, 375.0, 375.0, 385.0, 385.0, 395.0, 395.0, 400.0, 400.0, 400.0],
    'content': ['', '', 'Request Parameter Removed, Field type missing, Violation', 'Path Removed w/o Deprecation', '', 'Request Type Change,Removed param, Interface missing', 'Path Removed w/o Deprecation', 'Path Removed w/o Deprecation', 'New Required Request Property', 'Response Success State Removed, Violation', 'Field type changed']
}
df = pd.DataFrame(data)

# 标记content是否包含多个逗号分隔字段
df['has_multiple'] = df['content'].fillna('').str.contains(',')

# 统计满足条件的唯一api_spec_id数量
valid_groups = df[df['has_multiple']].groupby('label')['api_spec_id'].nunique()

# 整理目标label类别并补全计数
target_labels = {'Nan': None, 'minor': 'minor', 'patch': 'patch', 'major': 'major'}
final_counts = {k: valid_groups.get(v, 0) for k, v in target_labels.items()}

# 输出预期格式结果
for label, count in final_counts.items():
    print(f"`{label}`: {count}")

输出结果

`Nan`: 1
`minor`: 2
`patch`: 0
`major`: 0

内容的提问来源于stack exchange,提问作者Brie MerryWeather

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 00:00:18