二元分类数据标记重叠可视化:矩阵计算方法问询
计算非对称特征比例矩阵的方法
针对你描述的0/1特征矩阵需求,核心是计算行特征为1时,列特征为1的条件概率,具体实现步骤如下:
核心逻辑
对于矩阵中位置(行特征X, 列特征Y),数值计算公式为:
P(Y=1 | X=1) = 样本中X=1且Y=1的数量 / 样本中X=1的数量
对角线位置因为X=Y,结果恒为1。
Python代码实现
假设你的DataFrame为df,包含A、B、C三个0/1列:
import pandas as pd # 获取所有特征列名 features = df.columns.tolist() # 初始化结果矩阵,索引和列名均为特征名 result_matrix = pd.DataFrame(index=features, columns=features) for row_feature in features: # 筛选当前行特征为1的样本子集 row_subset = df[df[row_feature] == 1] # 处理行特征无1的情况(避免除以0) if len(row_subset) == 0: result_matrix.loc[row_feature] = 0.0 continue # 遍历列特征计算条件比例 for col_feature in features: if row_feature == col_feature: # 对角线固定为1 result_matrix.loc[row_feature, col_feature] = 1.0 else: # 计算X=1且Y=1的样本数占X=1样本数的比例 count_both = len(row_subset[row_subset[col_feature] == 1]) result_matrix.loc[row_feature, col_feature] = count_both / len(row_subset) # 转换为数值类型确保格式正确 result_matrix = result_matrix.astype(float)
示例验证
假设你的df有如下数据:
| A | B | C |
|---|---|---|
| 1 | 1 | 0 |
| 1 | 0 | 1 |
| 0 | 1 | 1 |
| 1 | 1 | 1 |
计算后得到的矩阵为:
| A | B | C | |
|---|---|---|---|
| A | 1.0 | 0.67 | 0.67 |
| B | 0.5 | 1.0 | 1.0 |
| C | 0.5 | 0.5 | 1.0 |
比如A-B的值是A=1的3个样本中B=1的有2个,所以2/3≈0.67;B-A是B=1的2个样本中A=1的有1个,所以1/2=0.5,完全符合你描述的非对称特性。
内容的提问来源于stack exchange,提问作者coder_bob
相关产品推荐
相关产品推荐

