You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何移除XY散点图中的异常值?附Python散点图绘制参考代码

可选解决方案

1. 基于2D直方图的密度过滤

你现有绘图代码已使用二维直方图统计密度,可直接复用该逻辑过滤低密度区域的离群点,无需新增依赖:

# 接你原有DataFrame生成逻辑
x = df['Data1'].values
y = df['Data2'].values

# 复用你现有直方图参数bins=45,统计每个网格的样本数
counts, x_edges, y_edges = np.histogram2d(x, y, bins=45)
# 密度阈值可调整,此处取所有非空网格计数的20分位数,分位数越高过滤越强
density_threshold = np.percentile(counts[counts > 0], 20)

# 匹配每个点所属的网格位置
x_idx = np.digitize(x, x_edges) - 1
y_idx = np.digitize(y, y_edges) - 1
# 仅保留高密度网格内的点
filter_mask = counts[x_idx, y_idx] > density_threshold
x_filtered = x[filter_mask]
y_filtered = y[filter_mask]

# 后续使用x_filtered、y_filtered绘图即可

2. DBSCAN密度聚类过滤

如果需要更精准的密度判别,使用基于密度的DBSCAN算法,自动将中间孤立点标记为噪声移除:

from sklearn.cluster import DBSCAN
from sklearn.preprocessing import StandardScaler

# 标准化数据消除量纲影响
X = StandardScaler().fit_transform(df[['Data1', 'Data2']])
# eps为邻域距离阈值,min_samples为邻域最少样本数,可根据实际效果调整
dbscan = DBSCAN(eps=0.05, min_samples=10).fit(X)
# 标签为-1的是噪声点,直接过滤
filter_mask = dbscan.labels_ != -1
x_filtered = df['Data1'].values[filter_mask]
y_filtered = df['Data2'].values[filter_mask]

3. 几何边界过滤(已知红线参数时使用)

如果图中两条红线的函数表达式已知,直接通过距离判断保留红线附近的点,效率最高:

# 替换为你实际的两条红线直线参数 y = a*x + b
a1, b1 = 1.2, 0.5
a2, b2 = 1.2, -0.5
# 距离阈值可调整,值越小保留的点越靠近红线
dist_threshold = 0.1

# 计算每个点到两条直线的垂直距离
dist1 = np.abs(a1 * x - y + b1) / np.sqrt(a1**2 + 1)
dist2 = np.abs(a2 * x - y + b2) / np.sqrt(a2**2 + 1)
# 保留靠近任意一条红线的点
filter_mask = (dist1 < dist_threshold) | (dist2 < dist_threshold)
x_filtered = x[filter_mask]
y_filtered = y[filter_mask]

内容的提问来源于stack exchange,提问作者Burak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 04:45:04