You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python合并CSV不同类型列生成2D直方图时如何避免NaN值

问题

我正在将两类CSV数据(dense和collision)中的列进行处理,以生成2D直方图。每类数据均包含type列,其中type=0代表big类型,type=1代表small类型。单独绘制big或small类型的a列与|f|列的2D直方图时一切正常,但将big和small的a、|f|列直接相加后绘制,结果出现了近90%的NaN值(原始数据无NaN),导致直方图异常。

数据示例(collision类型)

TIMESTEP id     type    a      |f|     |v|  
20000   4737     0     9.81  1.31495  4.18007   
40000  11991     1     9.81  4.43794  4.17909   
50000  15725     1     9.81  4.43794  4.17810     
30000   8209     0     9.81  4.43794  4.17810     
15000   3545     0     9.81  1.31495  4.17810   
30000   8269     0     9.81  4.43794  4.17810    
10000   2077     1     9.81  1.31495  4.17712   
20000   5079     0     9.81  1.31495  4.17712   

相关代码

from cProfile import label
from matplotlib.colors import LogNorm
import matplotlib.pyplot as plt
import numpy as np

df_collision_big = df_collision[df_collision['type'] == 0]
df_collision_small = df_collision[df_collision['type'] == 1]

df_dense_big = df_dense[df_dense['type'] == 0]
df_dense_small = df_dense[df_dense['type'] == 1]


plt.subplots(figsize=(14, 6))
# 调整子图间距
plt.subplots_adjust(wspace=0.5, hspace=0.6)
plt.subplot(231)
plt.hist2d(df_collision_small['a'], df_collision_small['|f|'], bins=np.linspace(0,70,15), norm=LogNorm())
plt.colorbar()
plt.xlabel('a')
plt.ylabel('|f|')
plt.title('Collision: Small')
plt.subplot(232)
plt.hist2d(df_collision_big['a'], df_collision_big['|f|'], bins=np.linspace(0,70,15), norm=LogNorm())
plt.colorbar()
plt.xlabel('a')
plt.ylabel('|f|')
plt.title('Collision: Big')
plt.subplot(233)
plt.hist2d(df_collision_big['a'] + df_collision_small['a'], df_collision_big['|f|'] + df_collision_small['|f|'], bins=np.linspace(0,70,15), norm=LogNorm())
plt.colorbar()
plt.xlabel('a')
plt.ylabel('|f|')
plt.title('Collision: Small + Big')

plt.subplot(234)
plt.hist2d(df_dense_small['a'], df_dense_small['|f|'], bins=np.linspace(0,70,15), norm=LogNorm())
plt.colorbar()
plt.xlabel('a')
plt.ylabel('|f|')
plt.title('Dense: Small')
plt.subplot(235)
plt.hist2d(df_dense_big['a'], df_dense_big['|f|'], bins=np.linspace(0,70,15), norm=LogNorm())
plt.colorbar()
plt.xlabel('a')
plt.ylabel('|f|')
plt.title('Dense: Big')
plt.subplot(236)
plt.hist2d(df_dense_big['a'] + df_dense_small['a'], df_dense_big['|f|'] + df_dense_small['|f|'], bins=np.linspace(0,70,15), norm=LogNorm())
plt.colorbar()
plt.xlabel('a')
plt.ylabel('|f|')
plt.title('Dense: Big + Small')
plt.savefig('hist2d.png', dpi=300)
plt.show()

经统计,df_collision_small['a']长度为13772,df_collision_big['a']长度为14060,相加后结果出现大量NaN,期望得到与Dense:Small+Big类似的正常直方图,寻求解决该问题的方法。

解决方案

问题原因

Pandas中对长度不一致的Series进行元素相加时,会根据索引自动对齐。如果两个Series的索引无重叠或长度不同,无法对齐的位置就会生成NaN。你这里df_collision_small和df_collision_big长度不同、索引无对应关系,直接相加自然会出现大量NaN。而Dense数据可能刚好两类样本长度相同、索引完全对齐,所以相加后无NaN,但这只是巧合,并非正确做法。

正确实现方式

要绘制big和small类型合并后的2D直方图,不需要将对应列相加,而是应该把两类数据的a和|f|样本合并到一起,再传入hist2d函数:

方法1:用NumPy合并数组

# 修改Collision: Small + Big 子图代码
plt.subplot(233)
# 合并small和big的a列数据
combined_collision_a = np.concatenate([df_collision_small['a'], df_collision_big['a']])
# 合并small和big的|f|列数据
combined_collision_f = np.concatenate([df_collision_small['|f|'], df_collision_big['|f|']])
plt.hist2d(combined_collision_a, combined_collision_f, bins=np.linspace(0,70,15), norm=LogNorm())
plt.colorbar()
plt.xlabel('a')
plt.ylabel('|f|')
plt.title('Collision: Small + Big')

# 修改Dense: Big + Small 子图代码
plt.subplot(236)
combined_dense_a = np.concatenate([df_dense_small['a'], df_dense_big['a']])
combined_dense_f = np.concatenate([df_dense_small['|f|'], df_dense_big['|f|']])
plt.hist2d(combined_dense_a, combined_dense_f, bins=np.linspace(0,70,15), norm=LogNorm())
plt.colorbar()
plt.xlabel('a')
plt.ylabel('|f|')
plt.title('Dense: Big + Small')

方法2:用Pandas合并DataFrame

也可以先合并两类样本的DataFrame,再提取列绘制:

# 合并Collision的small和big数据
df_collision_combined = pd.concat([df_collision_small, df_collision_big])
plt.subplot(233)
plt.hist2d(df_collision_combined['a'], df_collision_combined['|f|'], bins=np.linspace(0,70,15), norm=LogNorm())
plt.colorbar()
plt.xlabel('a')
plt.ylabel('|f|')
plt.title('Collision: Small + Big')

# Dense数据同理
df_dense_combined = pd.concat([df_dense_small, df_dense_big])
plt.subplot(236)
plt.hist2d(df_dense_combined['a'], df_dense_combined['|f|'], bins=np.linspace(0,70,15), norm=LogNorm())
plt.colorbar()
plt.xlabel('a')
plt.ylabel('|f|')
plt.title('Dense: Big + Small')

两种方法都能将所有small和big样本整合到一起,生成正常的2D直方图,不会出现NaN问题。


内容的提问来源于stack exchange,提问作者Codeholic

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 03:05:23