如何在Plotly Python直方图中设置日销售额超75的溢出分箱?
问题描述
我使用以下代码绘制直方图:
import matplotlib.pyplot as plt import plotly.express as px df = px.data.tips() fig = px.histogram(dataset, x="salesperday") fig.update_traces(xbins=dict( # bins used for histogram start=0.0, end=1400.0, size=40 )) fig.show()
请问如何自定义设置日销售额大于75的溢出分箱,实现类似Excel中的效果(将所有大于75的数值归为一个单独的分箱组)?
解决方案
要实现这种溢出分箱效果,你可以通过预处理数据或者自定义坐标轴刻度与分箱两种方式完成,以下是具体实现:
方法1:预处理数据(推荐)
先对原始数据的salesperday列进行分组,把大于75的数值统一标记为"75+",再用处理后的数据绘制直方图:
import plotly.express as px import pandas as pd # 替换为你的真实数据 df = pd.DataFrame({ "salesperday": [20, 35, 50, 65, 70, 75, 80, 90, 100, 110, 60, 45] }) # 预处理:将大于75的数值替换为"75+" df["sales_binned"] = df["salesperday"].apply(lambda x: x if x <=75 else "75+") # 绘制直方图 fig = px.histogram(df, x="sales_binned") fig.update_layout(xaxis_title="日销售额", yaxis_title="频数") fig.show()
方法2:自定义分箱与坐标轴刻度
如果不想修改原始数据,可以通过设置xbins的范围,再手动修改坐标轴的刻度文本,把超过75的分箱显示为"75+":
import plotly.express as px import pandas as pd # 替换为你的真实数据 df = pd.DataFrame({ "salesperday": [20, 35, 50, 65, 70, 75, 80, 90, 100, 110, 60, 45] }) # 绘制直方图,设置分箱上限为75 fig = px.histogram(df, x="salesperday") fig.update_traces( xbins=dict( start=0, end=75, size=15 # 自定义每个分箱的区间大小 ), # 可选:给溢出分箱设置不同颜色区分 marker_color=["#1f77b4"]*5 + ["#ff7f0e"] ) # 自定义坐标轴刻度,替换溢出分箱的显示文本 fig.update_layout( xaxis=dict( tickvals=[7.5, 22.5, 37.5, 52.5, 67.5, 95], ticktext=["0-15", "15-30", "30-45", "45-60", "60-75", "75+"], title="日销售额" ), yaxis_title="频数" ) fig.show()
注意事项
- 你的原始代码存在两处小问题:定义了
df但调用时用了dataset,需保持变量名一致;px.data.tips()数据集没有salesperday字段,需替换为包含该字段的真实数据或模拟数据 - 方法1更直观,适合需要直接展示分组标签的场景;方法2适合保留原始数据结构,仅在可视化层面调整的场景
内容的提问来源于stack exchange,提问作者python_pi
相关产品推荐
相关产品推荐

