使用plotly.express(color参数)搭配按钮时数据错乱问题排查
问题描述
当使用px.scatter(color='key')参数绘制散点图时,Plotly交互按钮表现异常:点击任意按钮后数据错乱,点击用于展示y轴petal_length的Plength按钮也无法回到初始图表;但不添加color参数时按钮功能正常。
功能正常的代码
from plotly import express as px import pandas as pd import seaborn as sns df = sns.load_dataset('iris') # df.head() # > sepal_length sepal_width petal_length petal_width species # > 0 5.1 3.5 1.4 0.2 setosa # > 1 4.9 3.0 1.4 0.2 setosa # > 2 4.7 3.2 1.3 0.2 setosa # > 3 4.6 3.1 1.5 0.2 setosa # > 4 5.0 3.6 1.4 0.2 setosa # df.species.unique() # > array(['setosa', 'versicolor', 'virginica'], dtype=object) fig = px.scatter( df, x=df.index, y="petal_length", #color="species", title='Iris', ) updatemenus = [ dict( type="buttons", buttons=[ dict( label="Plength", method="update", args=[ { "y": [df['petal_length']] }, { "title": "Petal length vs index", "yaxis": {"title": "petal_length"}, }, ], ), dict( label="Slength", method="update", args=[ { "y": [df['sepal_length']] }, { "title": "Sepal length vs index", "yaxis": {"title": "sepal_length"}, }, ], ), ], ) ] # Update layout with the updatemenu fig.update_layout(updatemenus=updatemenus)
按钮点击后数据错乱的代码
from plotly import express as px import pandas as pd import seaborn as sns df = sns.load_dataset('iris') df.index fig = px.scatter( df, x=df.index, y="petal_length", color="species", title='Iris', ) updatemenus = [ dict( type="buttons", buttons=[ dict( label="Plength", method="update", args=[ { "y": [df['petal_length']] }, { "title": "Petal length vs index", "yaxis": {"title": "petal_length"}, }, ], ), dict( label="Slength", method="update", args=[ { "y": [df['sepal_length']] }, { "title": "Sepal length vs index", "yaxis": {"title": "sepal_length"}, }, ], ), ], ) ] # Update layout with the updatemenu fig.update_layout(updatemenus=updatemenus)
异常原因
核心问题在于使用color="species"时,Plotly Express会自动将数据按species分组,生成多个独立的trace(每个物种对应一个trace);而未加color参数时,整个数据集是一个单独的trace。
你的按钮更新逻辑中,args里的y只传入了一个完整的列数组(比如[df['petal_length']]),这会导致:
- 点击按钮时,Plotly会把所有现有trace的y值都替换成这个完整数组,破坏了原来的分组结构,造成数据错乱;
- 点击Plength按钮时,无法恢复到初始的分组trace状态,因为传入的y值是未分组的整体数据,和初始的多trace结构不匹配。
要修复这个问题,需要在按钮的args中,按分组分别传入对应的数据,保持trace数量一致。比如Plength按钮的y参数应该写成:
"y": [ df[df.species == 'setosa']['petal_length'], df[df.species == 'versicolor']['petal_length'], df[df.species == 'virginica']['petal_length'] ]
Slength按钮同理,传入对应分组的sepal_length数据,这样就能匹配初始的多trace结构,按钮功能恢复正常。
内容的提问来源于stack exchange,提问作者Matthias Arras
相关产品推荐
相关产品推荐

