直方图数据范围自动检测为[nan,nan],如何丢弃NaN值解决报错?
Hey there! Let's break down the two questions you have about the ValueError when plotting a histogram with seaborn, plus fix the issue.
1. 为什么直方图没有自动忽略 NaN 值?
The root cause here lies in how numpy's histogram calculation works under the hood. When you pass an array with NaN values to np.histogram() (which seaborn/matplotlib uses to generate histograms), it tries to auto-detect the data range using np.min() and np.max(). But if there's even one NaN in the array, both np.min() and np.max() will return NaN—resulting in the invalid [nan, nan] range that triggers your error.
On top of that, you're using the fit=scipy.stats.norm parameter to fit a normal distribution. This fitting process can't handle NaN values either, so even if the histogram could skip NaNs, the fitting step would fail anyway and halt the entire plotting process.
2. 手动丢弃 NaN 值的解决方案
The easiest fix is to filter out NaN values using pandas' dropna() method. Here's how to adjust your code:
import seaborn as sns import scipy.stats # First, clean the data by removing NaNs clean_response = g['resp'].dropna() # Define your parameters as before par = {'hist': True, 'kde': False, 'fit': scipy.stats.norm, 'bins': 'auto'} # Plot with the cleaned data sns.distplot(clean_response, color='blue', **par)
Alternatively, you can chain the dropna() call directly in the distplot line to keep it concise:
sns.distplot(g['resp'].dropna(), color='blue', **par)
Once you pass only valid, non-NaN data to distplot, numpy will correctly calculate the data range, the normal distribution fit will work as expected, and your histogram will render without errors.
内容的提问来源于stack exchange,提问作者WoJ

