如何在散点图绘制中为两列数据的前1%值设置不同颜色
解决方案
核心逻辑说明
你原有代码的循环会覆盖阈值存储变量,无法同时保留两列的分位阈值,我们可以先统一计算两列阈值,再为数据点打分类标签,最后按标签映射颜色绘图即可。
完整实现代码
import pandas as pd import matplotlib.pyplot as plt # 1. 读取数据 features = ['column1' , 'column2'] df = pd.read_csv('XX.csv', usecols=features, sep=';', encoding='ISO-8859-1') # 2. 计算两列的99%分位阈值 thresholds = df[features].quantile(q=0.99, numeric_only=True) col1_thres = thresholds['column1'] col2_thres = thresholds['column2'] # 3. 为数据点打分类标签(可根据需求调整分类规则) # 规则示例:0=普通点,1=仅column1前1%,2=仅column2前1%,3=两列都在前1% df['point_type'] = 0 df.loc[df['column1'] >= col1_thres, 'point_type'] += 1 df.loc[df['column2'] >= col2_thres, 'point_type'] += 2 # 4. 配置颜色映射,可自行修改颜色值 color_map = { 0: '#1f77b4', # 普通点:蓝色 1: '#ff7f0e', # 仅column1前1%:橙色 2: '#2ca02c', # 仅column2前1%:绿色 3: '#d62728' # 两列都在前1%:红色 } df['color'] = df['point_type'].map(color_map) # 5. 绘制散点图 plt.figure(figsize=(10,6)) # 以时间序列索引为x轴的双列散点图 plt.scatter(df.index, df['column1'], color=df['color'], label='column1', alpha=0.7) plt.scatter(df.index, df['column2'], color=df['color'], label='column2', alpha=0.7) # 如果需要绘制column1为x轴、column2为y轴的散点图,把上面两行替换为下面这行即可 # plt.scatter(df['column1'], df['column2'], color=df['color'], alpha=0.7) # 补充图表配置 plt.xlabel('时间序列索引') plt.ylabel('数值') plt.legend() plt.title('双列时间序列散点图(前1%数值特殊标记)') plt.show()
自定义调整说明
- 如果不需要区分是哪一列的前1%,只需要统一标记所有前1%的点,步骤3可以简化为:
df['is_top1'] = (df['column1'] >= col1_thres) | (df['column2'] >= col2_thres),然后颜色映射只需要两类即可。 - 散点的大小、透明度、颜色值都可以根据可视化需求自行调整。
内容的提问来源于stack exchange,提问作者bazinga
相关产品推荐
相关产品推荐

