用Upset Plot替代韦恩图分析蛋白质组学数据遇数据适配及多DF绘图问题
用Upset Plot分析蛋白质组学肽段数据的解决方案
一、适配现有布尔型数据绘图
你的当前数据结构已满足Upset Plot的要求,无需额外转换。以下以upsetplot库为例演示:
- 安装依赖(未安装时执行):
pip install upsetplot pandas
- 绘图代码:
import pandas as pd from upsetplot import plot # 调用你的现有数据 data = upset_df[['peptide','onlymq', 'onlysage', 'both']] # 统计各布尔组合对应的肽段数量 counts = data.groupby(['onlymq', 'onlysage', 'both']).size() # 生成Upset Plot plot(counts)
upsetplot会自动识别布尔列的组合,清晰展示仅MQ识别、仅Sage识别、两者共同识别的肽段分布,完全匹配你的分析目标。
二、直接输入多批次肽段列表绘图
完全支持像韦恩图一样,用多个肽段列表(DataFrame)直接生成Upset Plot,无需提前构建布尔列。步骤如下:
假设你有两个原始肽段DataFrame:
mq_peptides:仅含MQ识别的肽段列(列名如peptide)sage_peptides:仅含Sage识别的肽段列(列名如peptide)
代码示例:
import pandas as pd from upsetplot import from_contents, plot # 构建分组字典:键为分组名称,值为去重后的肽段集合 peptide_groups = { "MQ": set(mq_peptides['peptide'].drop_duplicates()), "Sage": set(sage_peptides['peptide'].drop_duplicates()) } # 转换为Upset Plot所需格式 upset_data = from_contents(peptide_groups) # 生成绘图 plot(upset_data)
from_contents方法会自动计算各分组的独有肽段、交集肽段数量,直接生成符合需求的Upset Plot,逻辑和韦恩图输入完全一致,更适配原始数据流程。
三、关键注意事项
- 若肽段存在重复值,需先执行去重:
df['peptide'].drop_duplicates() - 如需交互式图表,可替换为
plotly库实现,核心数据处理逻辑一致
内容的提问来源于stack exchange,提问作者user23980869
相关产品推荐
相关产品推荐

