解决Pandas分组计数时‘incompatible index’等报错问题
解决Pandas分组统计编码值计数的问题
问题背景
现有包含snp_id、is_severe、encoding列的DataFrame,需要按snp_id和is_severe分组,统计每组中encoding列里one和two的数量,输出指定格式的结果。
之前尝试的代码出现索引不匹配、维度错误的问题:
# 第一次尝试 df.groupby(["snp_id","is_sever_f","encoding_str"])["encoding_str"].count() # 报错:incompatible index of inserted column with frame index # 第二次尝试 df["count"]=df.groupby(["snp_id","is_sever_f","encoding_str"],as_index=False)["encoding_str"].count() # 报错:Expected a 1D array, got an array with shape (2532831, 3)
正确实现方法
方法1:分组+值计数+透视
通过groupby分组后统计encoding的取值频率,再用unstack将行转列,补全缺失值为0:
# 分组统计encoding的取值计数 grouped_counts = df.groupby(["snp_id", "is_severe"])["encoding"].value_counts() # 将encoding的取值转为列,缺失值填充0 result = grouped_counts.unstack(fill_value=0) # 重命名列名并重置索引 result = result.rename(columns={"one": "encoding_one", "two": "encoding_two"}).reset_index()
方法2:使用交叉表pd.crosstab
直接用crosstab生成分组统计结果,更简洁:
import pandas as pd # 生成交叉表,行是snp_id+is_severe,列是encoding的取值 result = pd.crosstab([df["snp_id"], df["is_severe"]], df["encoding"]).reset_index() # 重命名列名匹配需求 result = result.rename(columns={"one": "encoding_one", "two": "encoding_two"})
错误原因说明
- 第一次代码将
encoding_str也加入分组,得到的是三级索引的Series,无法直接与原DataFrame的索引匹配,导致插入列时报错; - 第二次代码试图将分组返回的多列DataFrame赋值给原DataFrame的单个列,而单列只能接受一维数据,因此触发维度不匹配的错误。
内容的提问来源于stack exchange,提问作者Eliza Romanski
相关产品推荐
相关产品推荐

