如何为nd.array的百分位数图标注坐标及获取拐点上方数据
解决百分位数图拐点筛选与annotate标注问题
首先,我先梳理下你的需求:你想要绘制百分位数曲线,找到83%分位点对应的阈值(也就是该分位点的y坐标),然后筛选出原始数据中高于这个阈值的部分,同时解决当前annotate方法遇到的报错问题。
先看你代码里的几个核心问题:
pchange的长度远大于perc和p的长度,你用enumerate(pchange)去循环,然后调用perc[a]会直接触发索引越界错误——因为perc只有13个元素(和p的长度一致),但pchange有130+个元素。- 你混淆了百分位数计算后的数据对应关系:
mlab.prctile(pchange, p)返回的是p中每个百分位数对应的数值,所以perc的第i个元素对应p[i]百分位的数值。
接下来一步步解决你的问题:
1. 获取83%分位点的阈值
我们可以直接从计算好的perc中提取83%分位点对应的数值,也可以用更常用的numpy.percentile直接计算,结果和mlab.prctile一致:
import numpy as np import matplotlib.pyplot as plt from matplotlib import mlab # 你的原始数据 p = np.array([0,10,20,30,40,50,60,70,80,83,84,85,90]) pchange=np.array([1,2,3,6,5,8,9,7,4,5,6,9,8,5,2,3,6,4,25,36,14,65,98,98,54,25,26,23,24,27,28,26,24,262,1,156,31,51,351,651,35,153,135,1,5,31,68,3,5,61,354,685,16,813,51,685,681,35,68,135,1685,1354,135,415,135,153,413,513,56,513,213,651,354,51,35,135,135,135,438,535,468,53,8,35,4,648,468,535,468,46,8,498,498,749,8798,798,79,8798,7,979,879,8,97,9,79,7,9798,798,78,979,87,974,65,498,46,8,98,79,878,978,65,984,98,49,9,569,949,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888,888]) # 计算百分位数 perc = mlab.prctile(pchange, p) # 找到83%分位点对应的索引 idx_83 = np.where(p == 83)[0][0] # 获取对应的阈值 threshold_83 = perc[idx_83] print(f"83%分位点对应的阈值为: {threshold_83}") # 筛选出pchange中高于该阈值的数据 filtered_data = pchange[pchange > threshold_83] print(f"筛选出的高于83%分位点的数据: {filtered_data}")
2. 修正绘图与annotate标注问题
你的annotate错误是因为循环对象不对,应该循环perc和p的对应关系,而不是pchange。另外,绘图时的x轴坐标计算需要注意:(len(perc)-1)*p/100是把百分位数值映射到0到len(perc)-1的x轴位置,这样xticks能准确对应。
修正后的绘图代码:
plt.figure(figsize=(10,6)) # 绘制百分位数曲线 plt.plot(perc, label='Percentile Curve') # 绘制每个百分位点的散点 x_coords = (len(perc)-1)*p/100. plt.plot(x_coords, perc, 'ro', markersize=6, label='Percentile Points') # 设置x轴刻度 plt.xticks(x_coords, map(str, p)) # 标注每个百分位点的数值,调整位置避免重叠 for x, y, pc in zip(x_coords, perc, p): plt.annotate(f"{y:.1f}", (x, y), xytext=(x+0.1, y+5), fontsize=9) # 绘制83%分位点的阈值线,方便观察拐点 plt.axhline(y=threshold_83, color='r', linestyle='--', label=f'83% Percentile Threshold: {threshold_83:.1f}') plt.legend() plt.xlabel('Percentile') plt.ylabel('Value') plt.title('Percentile Plot with 83% Threshold') plt.show()
关键说明
- 如果你只是需要计算单个百分位数,推荐使用
numpy.percentile,用法更简洁:np.percentile(pchange, 83)直接得到83%分位点的数值。 - 筛选数据时,直接用布尔索引
pchange[pchange > threshold_83]就能快速得到所有高于阈值的数据。 annotate时,要确保循环的是x_coords、perc、p这三个长度一致的数组,同时用xytext调整标注位置,避免和散点重叠。
内容的提问来源于stack exchange,提问作者Bamwani
相关产品推荐
相关产品推荐

