You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何创建按两个DataFrame点数占比配色的2D密度图并解决像素化问题?

优化2D点数比值图的像素化问题

我有两个定义在(x,y)平面上的2D DataFrame——df1和df2,它们位于x,y平面的相似区域。我希望创建一张以df2点数除以df1点数为颜色编码的图像,但生成的图存在明显像素化问题,求优化建议。

我当前的代码

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt

df1 = pd.DataFrame(np.random.random((10000,2)), columns=list('AB'))
df2 = pd.DataFrame(np.random.random((10000,2)), columns=list('AB'))

upper_lim_A = 0.8
lower_lim_A = 0.3

upper_lim_B = 0.7
lower_lim_B = 0.2

fontsize = 8
N_bins = 10
hist_threshold = 1 # 分箱内点数低于此值视为无效

%matplotlib inline
# 生成二维直方图
hist_all, *edges_all = np.histogram2d(df1['B'], df1['A'], bins=N_bins, range = ((lower_lim_B, upper_lim_B), (lower_lim_A,upper_lim_A)))
hist_vars, *edges_vars = np.histogram2d(df2['B'], df2['A'], bins=N_bins, range = ((lower_lim_B, upper_lim_B), (lower_lim_A,upper_lim_A)))

fig, ax1 = plt.subplots( figsize=(10.5,2.5*3))
# 注:原代码中size_la未定义,此处替换为已定义的fontsize
ax1.tick_params(direction='out', length=6, width=2, colors='k', grid_alpha=1, labelsize=fontsize)

# 计算df2/df1的比值
ratio = hist_vars/hist_all
# 替换NaN为0
ratio[np.isnan(ratio)] = 0

# 备份原始比值
ratio_copy = np.copy(ratio)
# 低计数分箱设为NaN(隐藏无效区域)
ratio[hist_all < hist_threshold*0.9] = np.nan

color_style = 'inferno'
vmax = 1.1

im = ax1.imshow(ratio, alpha=1,origin='lower',cmap=color_style,vmin=0, aspect='auto', extent=[lower_lim_A,upper_lim_A,lower_lim_B, upper_lim_B])

from mpl_toolkits.axes_grid1 import make_axes_locatable

ax0 = fig.add_subplot(111)
ax0.set_frame_on(False)
ax0.xaxis.set_ticks([])
ax0.yaxis.set_ticks([])

# 创建颜色条容器
divider = make_axes_locatable(ax0)
cax = divider.append_axes("top", size="5%", pad=0.1)

# 添加颜色条
cbar = fig.colorbar(im, cax=cax, orientation='horizontal')
cbar.ax.tick_params(labelsize=fontsize)
cbar.set_label(r'$\mathrm{N}_{\mathrm{df2}}/\mathrm{N}_{\mathrm{df1}}$', fontsize=25)
cax.xaxis.set_ticks_position('top')
cax.xaxis.set_label_position('top')
ax1.invert_yaxis()

当前生成的图像:呈现明显的方块状像素,区域过渡生硬,细节丢失。

优化方案

1. 增加分箱数+插值平滑

像素化的核心原因是N_bins=10分箱太少,每个分箱尺寸过大。调整分箱数并启用插值即可快速改善:

  • 将N_bins调整为50-100(根据数据量灵活调整)
  • 在imshow中添加interpolation='bilinear'或'bicubic'参数,实现平滑过渡:
# 修改分箱数
N_bins = 50
# 绘图时添加插值
im = ax1.imshow(ratio, alpha=1, origin='lower', cmap=color_style, vmin=0, vmax=1.1, 
                aspect='auto', extent=[lower_lim_A,upper_lim_A,lower_lim_B, upper_lim_B], 
                interpolation='bilinear')

2. 改用核密度估计(KDE)替代直方图

如果数据量足够(如你的10000条数据),KDE能生成更平滑的密度比值图,完全避免分箱带来的像素感:

from scipy.stats import gaussian_kde

# 整理数据格式
xy1 = np.vstack([df1['A'], df1['B']])
xy2 = np.vstack([df2['A'], df2['B']])

# 计算两个数据集的KDE密度函数
kde1 = gaussian_kde(xy1)
kde2 = gaussian_kde(xy2)

# 创建高密度网格
xgrid = np.linspace(lower_lim_A, upper_lim_A, 200)
ygrid = np.linspace(lower_lim_B, upper_lim_B, 200)
X, Y = np.meshgrid(xgrid, ygrid)
xy_grid = np.vstack([X.ravel(), Y.ravel()])

# 计算密度比值
density1 = kde1(xy_grid).reshape(X.shape)
density2 = kde2(xy_grid).reshape(X.shape)
ratio_kde = density2 / density1

# 处理异常值
ratio_kde[np.isnan(ratio_kde)] = 0
ratio_kde[density1 < 1e-6] = np.nan  # 过滤密度极低的无效区域

# 用pcolormesh绘图
im = ax1.pcolormesh(X, Y, ratio_kde, cmap=color_style, vmin=0, vmax=1.1)

3. 替换绘图函数为pcolormesh

imshow默认会对矩阵做拉伸对齐,改用pcolormesh能更精准匹配分箱边缘,减少像素感:

# 获取分箱边缘
x_edges = edges_all[1]
y_edges = edges_all[0]
# 用pcolormesh绘制直方图比值
im = ax1.pcolormesh(x_edges, y_edges, ratio, cmap=color_style, vmin=0, vmax=1.1)

4. 低计数区域平滑处理

如果不想过滤低计数分箱,可以用邻域平均平滑这些区域:

from scipy.ndimage import uniform_filter

# 对原始比值做邻域平滑(size控制窗口大小)
smoothed_ratio = uniform_filter(ratio_copy, size=2)
# 保留低计数区域的NaN标记
smoothed_ratio[hist_all < hist_threshold*0.9] = np.nan
# 用平滑后的数据绘图
im = ax1.imshow(smoothed_ratio, alpha=1, origin='lower', cmap=color_style, vmin=0, vmax=1.1, 
                aspect='auto', extent=[lower_lim_A,upper_lim_A,lower_lim_B, upper_lim_B],
                interpolation='bilinear')

内容的提问来源于stack exchange,提问作者Cruz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 07:55:08