如何用Plotly Express实现分类变量散点无重叠展示?
在Plotly Express中复现JMP的分类交叉散点抖动效果
你需要对比两个数据库返回的同一项目的分类赋值,计划将两个数据库的分类分别作为坐标轴,使用px.scatter可视化交叉情况,但遇到问题:px.scatter没有内置抖动数据点避免重叠的选项,尝试scattermode='group'参数也无效。当前px.scatter的效果是所有数据点堆叠在一起,而你已在JMP中实现了能清晰展示每个交叉点数据点数量的效果,希望在Plotly Express中复现。
最小可复现代码(MWE)
import pandas as pd import plotly.express as px d = {'Document_Type_x': ['Research Article', 'Research Article', 'Letter to the Editor', 'Letter to the Editor', 'Letter to the Editor'], 'Document_Type_y': ['Article', 'Article', 'Letter', 'Letter', 'Letter']} df = pd.DataFrame(data=d) fig = px.scatter(df, x='Document_Type_x', y='Document_Type_y') fig.update_layout(scattermode='group', scattergap=.9) fig.update_xaxes(categoryorder = 'category ascending') fig.update_yaxes(categoryorder = 'category ascending') fig.show()
解决方案
方法1:手动添加随机抖动(最接近JMP效果)
Plotly Express没有内置抖动参数,但可以通过给分类变量的数值化结果添加微小随机偏移,让同组的点分散开:
import pandas as pd import plotly.express as px import numpy as np d = {'Document_Type_x': ['Research Article', 'Research Article', 'Letter to the Editor', 'Letter to the Editor', 'Letter to the Editor'], 'Document_Type_y': ['Article', 'Article', 'Letter', 'Letter', 'Letter']} df = pd.DataFrame(data=d) # 将分类转为数值编码,再添加随机抖动 df['x_jitter'] = pd.Categorical(df['Document_Type_x']).codes + np.random.uniform(-0.15, 0.15, size=len(df)) df['y_jitter'] = pd.Categorical(df['Document_Type_y']).codes + np.random.uniform(-0.15, 0.15, size=len(df)) # 基于抖动后的数值绘制散点图 fig = px.scatter(df, x='x_jitter', y='y_jitter') # 将坐标轴刻度替换为原始分类名称 fig.update_xaxes( tickvals=pd.Categorical(df['Document_Type_x']).codes, ticktext=pd.Categorical(df['Document_Type_x']).categories, categoryorder='category ascending' ) fig.update_yaxes( tickvals=pd.Categorical(df['Document_Type_y']).codes, ticktext=pd.Categorical(df['Document_Type_y']).categories, categoryorder='category ascending' ) # 调整轴范围,避免点超出可视区域 fig.update_layout( xaxis_range=[-0.5, max(pd.Categorical(df['Document_Type_x']).codes) + 0.5], yaxis_range=[-0.5, max(pd.Categorical(df['Document_Type_y']).codes) + 0.5] ) fig.show()
方法2:用气泡图展示交叉点计数(更直观)
如果不需要展示单个数据点,直接统计每个分类交叉组合的样本数量,用气泡大小体现数量:
import pandas as pd import plotly.express as px d = {'Document_Type_x': ['Research Article', 'Research Article', 'Letter to the Editor', 'Letter to the Editor', 'Letter to the Editor'], 'Document_Type_y': ['Article', 'Article', 'Letter', 'Letter', 'Letter']} df = pd.DataFrame(data=d) # 统计每个分类交叉组合的样本数 count_df = df.groupby(['Document_Type_x', 'Document_Type_y']).size().reset_index(name='Count') # 绘制气泡图,气泡大小对应样本数量 fig = px.scatter( count_df, x='Document_Type_x', y='Document_Type_y', size='Count', size_max=60 # 控制最大气泡尺寸 ) fig.update_xaxes(categoryorder='category ascending') fig.update_yaxes(categoryorder='category ascending') fig.show()
方法3:使用Plotly Graph Objects自定义散点
如果需要更精细的样式控制,可以直接用plotly.graph_objects实现抖动效果:
import pandas as pd import plotly.graph_objects as go import numpy as np d = {'Document_Type_x': ['Research Article', 'Research Article', 'Letter to the Editor', 'Letter to the Editor', 'Letter to the Editor'], 'Document_Type_y': ['Article', 'Article', 'Letter', 'Letter', 'Letter']} df = pd.DataFrame(data=d) # 将分类转为数值编码 x_cats = pd.Categorical(df['Document_Type_x']) y_cats = pd.Categorical(df['Document_Type_y']) fig = go.Figure() # 添加带抖动的散点 fig.add_trace(go.Scatter( x=x_cats.codes + np.random.uniform(-0.15, 0.15, len(df)), y=y_cats.codes + np.random.uniform(-0.15, 0.15, len(df)), mode='markers', marker=dict(size=10, color='#1f77b4') )) # 设置坐标轴刻度为原始分类名称 fig.update_xaxes( tickvals=x_cats.codes, ticktext=x_cats.categories, categoryorder='category ascending' ) fig.update_yaxes( tickvals=y_cats.codes, ticktext=y_cats.categories, categoryorder='category ascending' ) # 调整轴范围 fig.update_layout( xaxis_range=[-0.5, max(x_cats.codes) + 0.5], yaxis_range=[-0.5, max(y_cats.codes) + 0.5] ) fig.show()
内容的提问来源于stack exchange,提问作者eschares
相关产品推荐
相关产品推荐

