Seaborn中barplot与regplot置信区间结果不稳定原因咨询
问题原因分析:Seaborn绘图置信区间每次结果不一致
环境信息
- 系统:Linux Mint 21
- 开发环境:Spyder 5.3.3(Jupyter Notebook模式)
- Python版本:3.9.16(64位)
- 使用库:Pandas、GeoPandas、seaborn、matplotlib
操作步骤
- 导入Pandas、GeoPandas、seaborn、matplotlib等库
- 将gpkg文件导入为GeoDataFrame(gdf)
- 通过参数从原gdf提取数据生成多个新gdf,并存入DataFrame用于后续循环分析
- 整理两个gdf的关键信息为长格式DataFrame,结合seaborn boxplot与barplot绘图,用barplot展示箱线图上方的95%置信区间,代码如下:
### add CI ax2= sns.barplot( data=df_plot, x='bez', y="distance_overtaker", #palette="Blues", order = pos2, alpha=0.0, capsize=.1, n_boot=1000, errorbar=('ci', 95), errcolor= 'red', #errcolor='.26' = errwidth=0.7, ax=ax)
- 直接使用两个gdf,通过seaborn regplot绘图,代码如下:
for i, row in liste.iterrows(): sns.regplot(x=krit, y='distance_overtaker', data=row['gdf'], fit_reg=True, x_jitter=0, ci=95, ax=ax, scatter_kws={'alpha':1., 's':1}, label=row['bez'])
问题现象
- 执行第4步时,每次运行barplot的误差线(尤其是右侧)都可能不同,无明显规律
- 执行第5步时,每次运行regplot的置信区间带均有细微差异,呈现多种变化形式
原因解析
这是因为Seaborn计算置信区间时默认使用自助法(Bootstrapping),这种方法依赖随机抽样估算置信区间:
- 你的
barplot代码显式设置了n_boot=1000,表示用1000次自助抽样计算CI;regplot默认也会使用1000次自助抽样 - 每次运行代码时,自助抽样的随机样本都不相同,因此计算出的置信区间会存在细微波动
- 若要固定结果,可在代码开头设置全局随机种子,让每次抽样的随机序列完全一致,示例代码:
import numpy as np import random np.random.seed(42) random.seed(42)
内容的提问来源于stack exchange,提问作者Drödmbüddel
相关产品推荐
相关产品推荐

