You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计DataFrame中仅单站点及双站点检测到的个体数量?

统计单站点/双站点检测的个体数量

核心思路

先按个体ID分组,统计每个ID对应的唯一检测站点集合,再根据集合的长度和内容,分别统计仅单站点出现、双站点都出现的个体数量。

实现代码

1. 构造数据与导入依赖

import pandas as pd

# 构建用户提供的数据集
data = {
    'ID': [5994, 5994, 6017, 6018, 6020, 6020, 6021, 46507, 46508, 46511, 46512, 46512],
    'Location': ['Greshornish', 'Snizort', 'Greshornish', 'Greshornish', 'Greshornish', 'Snizort', 'Greshornish', 'Greshornish', 'Snizort', 'Greshornish', 'Greshornish', 'Snizort']
}
df = pd.DataFrame(data)

2. 分组统计每个ID的站点信息

# 按ID分组,获取每个ID对应的唯一站点集合
id_loc_groups = df.groupby('ID')['Location'].unique()

3. 分类统计目标数量

# 分别统计三类个体的数量
counts = {
    'Greshornish': sum(id_loc_groups.apply(lambda x: len(x) == 1 and x[0] == 'Greshornish')),
    'Snizort': sum(id_loc_groups.apply(lambda x: len(x) == 1 and x[0] == 'Snizort')),
    'Both': sum(id_loc_groups.apply(lambda x: len(x) == 2))
}

# 转换为预期格式的DataFrame
result_df = pd.DataFrame(list(counts.items()), columns=['Location', 'Count'])

4. 查看结果

执行print(result_df)后,输出与预期完全一致:

Location  Count
0  Greshornish      5
1      Snizort      1
2         Both      3

另一种更直观的实现方式

如果需要更清晰的中间统计结果,可以先聚合每个ID的站点数量和集合:

# 聚合每个ID的站点数量与站点集合
id_stats = df.groupby('ID').agg(
    site_count=('Location', 'nunique'),
    site_set=('Location', set)
)

# 筛选统计各类别
only_gresh = len(id_stats[(id_stats['site_count'] == 1) & (id_stats['site_set'] == {'Greshornish'})])
only_sniz = len(id_stats[(id_stats['site_count'] == 1) & (id_stats['site_set'] == {'Snizort'})])
both_sites = len(id_stats[id_stats['site_count'] == 2])

# 生成结果DataFrame
result_df = pd.DataFrame([
    ['Greshornish', only_gresh],
    ['Snizort', only_sniz],
    ['Both', both_sites]
], columns=['Location', 'Count'])

这种方式把中间统计指标明确展示,便于调试和扩展。

内容的提问来源于stack exchange,提问作者mikejwilliamson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 13:45:33