如何统计DataFrame中仅单站点及双站点检测到的个体数量?
统计单站点/双站点检测的个体数量
核心思路
先按个体ID分组,统计每个ID对应的唯一检测站点集合,再根据集合的长度和内容,分别统计仅单站点出现、双站点都出现的个体数量。
实现代码
1. 构造数据与导入依赖
import pandas as pd # 构建用户提供的数据集 data = { 'ID': [5994, 5994, 6017, 6018, 6020, 6020, 6021, 46507, 46508, 46511, 46512, 46512], 'Location': ['Greshornish', 'Snizort', 'Greshornish', 'Greshornish', 'Greshornish', 'Snizort', 'Greshornish', 'Greshornish', 'Snizort', 'Greshornish', 'Greshornish', 'Snizort'] } df = pd.DataFrame(data)
2. 分组统计每个ID的站点信息
# 按ID分组,获取每个ID对应的唯一站点集合 id_loc_groups = df.groupby('ID')['Location'].unique()
3. 分类统计目标数量
# 分别统计三类个体的数量 counts = { 'Greshornish': sum(id_loc_groups.apply(lambda x: len(x) == 1 and x[0] == 'Greshornish')), 'Snizort': sum(id_loc_groups.apply(lambda x: len(x) == 1 and x[0] == 'Snizort')), 'Both': sum(id_loc_groups.apply(lambda x: len(x) == 2)) } # 转换为预期格式的DataFrame result_df = pd.DataFrame(list(counts.items()), columns=['Location', 'Count'])
4. 查看结果
执行print(result_df)后,输出与预期完全一致:
Location Count 0 Greshornish 5 1 Snizort 1 2 Both 3
另一种更直观的实现方式
如果需要更清晰的中间统计结果,可以先聚合每个ID的站点数量和集合:
# 聚合每个ID的站点数量与站点集合 id_stats = df.groupby('ID').agg( site_count=('Location', 'nunique'), site_set=('Location', set) ) # 筛选统计各类别 only_gresh = len(id_stats[(id_stats['site_count'] == 1) & (id_stats['site_set'] == {'Greshornish'})]) only_sniz = len(id_stats[(id_stats['site_count'] == 1) & (id_stats['site_set'] == {'Snizort'})]) both_sites = len(id_stats[id_stats['site_count'] == 2]) # 生成结果DataFrame result_df = pd.DataFrame([ ['Greshornish', only_gresh], ['Snizort', only_sniz], ['Both', both_sites] ], columns=['Location', 'Count'])
这种方式把中间统计指标明确展示,便于调试和扩展。
内容的提问来源于stack exchange,提问作者mikejwilliamson
相关产品推荐
相关产品推荐

