You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Streamlit中ff.create_distplot运行过慢报错的解决方法咨询

Streamlit中Plotly distplot耗时久且报错的解决方法

我在基于Streamlit构建数据应用时,用50行的回归数据集(shape为(50,5))运行distplot代码,出现耗时极久且报错的问题,代码如下:

import streamlit as st
import pandas as pd
import plotly.figure_factory as ff
data =pd.read_csv("https://raw.githubusercontent.com/krishnaik06/Multiple-Linear-Regression/master/50_Startups.csv")
print(data.head())
print(data.shape)
hist_data =[data['R&D Spend'].values,data['Administration'].values,data['Marketing Spend'].values,data['Profit'].values]
groups =["R&D spend","Administration","Marketing Spend","Profit"]
fig =ff.create_distplot(hist_data=hist_data,group_labels=groups,bin_size=[0.1,0.25,0.5,0.75])
st.plotly_chart(fig,use_container_width=True)

报错截图如下:
报错截图

问题原因

报错核心是bin_size设置过小。数据集里的R&D Spend、Profit等列数值都是几万到几十万的量级,你设置的bin_size=[0.1,0.25,0.5,0.75]意味着要把每个变量分成上百万个区间,哪怕只有50条数据,这种计算量也会直接撑爆内存,导致耗时久且报错。

解决方法

调整bin_size为与数据量级匹配的数值,或者直接让Plotly自动计算合适的bin大小:

方案1:移除bin_size参数,让系统自动适配

import streamlit as st
import pandas as pd
import plotly.figure_factory as ff

data = pd.read_csv("https://raw.githubusercontent.com/krishnaik06/Multiple-Linear-Regression/master/50_Startups.csv")
hist_data = [data['R&D Spend'].values, data['Administration'].values, data['Marketing Spend'].values, data['Profit'].values]
groups = ["R&D spend", "Administration", "Marketing Spend", "Profit"]

# 不指定bin_size,Plotly会自动计算合理的区间
fig = ff.create_distplot(hist_data=hist_data, group_labels=groups)
st.plotly_chart(fig, use_container_width=True)

方案2:手动设置匹配数据量级的bin_size

根据各列的数值范围,设置合适的区间大小,比如:

import streamlit as st
import pandas as pd
import plotly.figure_factory as ff

data = pd.read_csv("https://raw.githubusercontent.com/krishnaik06/Multiple-Linear-Regression/master/50_Startups.csv")
hist_data = [data['R&D Spend'].values, data['Administration'].values, data['Marketing Spend'].values, data['Profit'].values]
groups = ["R&D spend", "Administration", "Marketing Spend", "Profit"]

# 根据数据量级设置bin_size,比如R&D Spend用5000为区间,Marketing Spend用10000为区间
fig = ff.create_distplot(hist_data=hist_data, group_labels=groups, bin_size=[5000, 5000, 10000, 10000])
st.plotly_chart(fig, use_container_width=True)

这两种方法都能避免因bin_size过小导致的计算过载问题,快速生成正常的分布图。

内容的提问来源于stack exchange,提问作者AI ML

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 17:23:10