You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为representative points DataFrame添加对应簇的点数统计列?

解决方法

方法1:使用value_counts() + map

这种方法简洁高效,适合簇ID与代表点DataFrame的索引直接对应的场景:

# 统计每个簇包含的点数
cluster_counts = points_df['cluster'].value_counts()

# 为代表点DataFrame添加点数列,匹配索引对应的簇计数,无数据的簇填充0并转为整数
representative_points['point_count'] = representative_points.index.map(cluster_counts).fillna(0).astype(int)

方法2:使用groupby() + merge

如果需要更灵活的匹配逻辑(比如簇ID不是代表点的索引,而是单独的列),可以用分组统计后合并的方式:

# 按cluster分组统计点数,将簇ID转为列以便合并
cluster_counts = points_df.groupby('cluster').size().reset_index(name='point_count')

# 左合并到代表点DataFrame,保留所有簇的行,无数据的簇填充0并整理格式
representative_points = representative_points.merge(
    cluster_counts,
    left_index=True,  # 用代表点的索引(簇ID)匹配
    right_on='cluster',
    how='left'
).fillna(0).astype({'point_count': int}).drop(columns='cluster')

注意事项

  • 如果你的代表点DataFrame中不是用索引存储簇ID,而是有单独的cluster列,只需把left_index=True改为left_on='cluster'即可。
  • fillna(0)是为了处理那些没有对应数据点的簇,确保结果中不会出现空值。

内容的提问来源于stack exchange,提问作者Stackoverflow Stackoverflow

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 08:56:05