如何用Python基于分数数据统计并绘制维恩图?
解决思路与代码实现
数据准备
你的数据集如下:
df = id testA testB 1 3 NA 1 1 3 2 2 NA 2 NA 1 2 0 0 3 NA NA 3 1 1
需求说明
需要统计并绘制维恩图,展示三类记录的数量:
- 同时有testA和testB有效值(非NA)的记录
- 仅testA有有效值、testB为NA的记录
- 仅testB有有效值、testA为NA的记录
预期统计结果:
Both tests: 3
A but not B: 2
B but not A: 1
R语言实现方案
1. 统计分组数量
# 加载数据 df <- data.frame( id = c(1,1,2,2,2,3,3), testA = c(3,1,2,NA,0,NA,1), testB = c(NA,3,NA,1,0,NA,1) ) # 标记分组 df$group <- case_when( !is.na(df$testA) & !is.na(df$testB) ~ "Both", !is.na(df$testA) & is.na(df$testB) ~ "A only", is.na(df$testA) & !is.na(df$testB) ~ "B only", TRUE ~ "Neither" ) # 统计各组数量 group_counts <- table(df$group) group_counts
运行输出:
A only B only Both Neither 2 1 3 1
2. 绘制维恩图
使用VennDiagram包生成可视化:
library(VennDiagram) # 提取目标数值 a_only <- group_counts["A only"] b_only <- group_counts["B only"] both <- group_counts["Both"] # 生成维恩图 venn.diagram( x = list(TestA = c(rep("A", a_only + both)), TestB = c(rep("B", b_only + both))), filename = NULL, fill = c("#4285F4", "#EA4335"), alpha = 0.5, main = "TestA vs TestB 记录分布", cat.cex = 1.2, cex = 1.5 ) grid.draw(last_plot())
Python语言实现方案
1. 统计分组数量
import pandas as pd # 加载数据 df = pd.DataFrame({ "id": [1,1,2,2,2,3,3], "testA": [3,1,2,None,0,None,1], "testB": [None,3,None,1,0,None,1] }) # 统计各组 both = len(df[(df['testA'].notna()) & (df['testB'].notna())]) a_only = len(df[(df['testA'].notna()) & (df['testB'].isna())]) b_only = len(df[(df['testA'].isna()) & (df['testB'].notna())]) # 打印预期结果 print(f"Both tests: {both}") print(f"A but not B: {a_only}") print(f"B but not A: {b_only}")
2. 绘制维恩图
使用matplotlib_venn包生成可视化:
from matplotlib_venn import venn2 import matplotlib.pyplot as plt # 生成维恩图 venn2(subsets=(a_only, b_only, both), set_labels=('TestA', 'TestB')) plt.title("TestA vs TestB 记录分布") plt.show()
内容的提问来源于stack exchange,提问作者Economist_Ayahuasca
相关产品推荐
相关产品推荐

