调用transform函数触发np.isnan TypeError:输入类型不支持
我正尝试改编一篇近期发表论文中的代码用于自有数据集,对应的.ipynb文件可从原项目获取。我已按照作者描述的R预处理流程格式化数据,这是我首次使用Python,若表述不清或有低级错误请见谅。
最初的问题出在读取鱼类参考分类数据集(原代码针对哺乳动物),我通过读取两个独立CSV文件(一个存参考峰值,一个存分类信息)并合并,得到包含Order、Family、Common Name、Species及肽段位置列的reference数据框,前4列为字符型,后续列为数值或NaN。
相关代码如下:
df = pd.read_csv("/content/drive/MyDrive/Archaeology 2024/Thesis Re-runs/PP_Combo_1/Combo_1_C66.csv", sep=',') ref_taxa = pd.read_csv('/content/drive/MyDrive/Archaeology 2024/Thesis Re-runs/Taxa.csv', sep=",", encoding='UTF-8', header=None, names=fields) ref_peaks = pd.read_csv('/content/drive/MyDrive/Archaeology 2024/Thesis Re-runs/Edited_ZooMS_Values.csv', sep=",", encoding='UTF-8', header=None, names=fields2) ref_taxa = ref_taxa.iloc[1:] ref_peaks = ref_peaks.iloc[1:] references = pd.concat([ref_taxa,ref_peaks], axis=1)
我要运行的transform函数代码如下:
def transform(df): df_copy = df.copy() # making a copy so I dont have to reload the dataframe everytime #then changes are made to df_copy x = np.arange(500, 3500, 0.1) pi = np.array(ref_peaks.iloc[0].to_numpy()) pi = pi[~np.isnan(pi)] resampled = np.zeros_like(x) for p in df_copy['mass']: resampled[np.logical_and(p-x>-0.3,p-x<1.3)] = 1 resampled_smooth = gaussian_filter1d(resampled, 100) plt.plot(x, resampled, label='sample', linewidth=1) plt.plot(x, resampled_smooth, label='sample', linewidth=1) n_spec = ref_peaks.shape[0] corrs = np.zeros(n_spec) for i in range(n_spec): y = np.zeros_like(x) pi = np.array(ref_peaks.iloc[i].to_numpy()) pi = pi[~np.isnan(pi)] for p in pi: y[np.logical_and(p-x>-0.3,p-x<1.3)] = 1 corrs[i] = np.corrcoef(y, resampled)[0,1] ranks = pd.DataFrame({'Order':reference['Order'][1:],'Family':reference['Family'][1:], 'Species':reference['Species'][1:], 'ZooMS_taxon':reference['ZooMS_taxon'][1:], 'Correlation':corrs}) idx = ranks.sort_values(by='Correlation', ascending=False).index[:6] - 1 # index [6] because the first 6 matches of species will be determined -- number depends on how many of the first n most likely species you want to look at all_d = pd.DataFrame(columns=ref_peaks.columns) for i in idx: y = np.zeros_like(x) pi = np.array(ref_peaks.iloc[i].to_numpy()) pi = pi[~np.isnan(pi)] for p in pi: y[np.logical_and(p-x>-0.3,p-x<1.3)] = 1 plt.figure() plt.plot(x, y * resampled) plt.title(reference['Species'][i+1]); v, ix, iy = np.intersect1d((np.round(pi, decimals=1)*10).astype(int), (x[y*resampled==1]*10).astype(int), return_indices=True) d = pd.DataFrame({k:p for k, p in zip(ref_peaks.columns[ix],v/10)}, index=pd.MultiIndex.from_tuples([(reference['Species'][i+1], reference['Family'][i+1], reference['Order'][i+1], reference['ZooMS_taxon'][i+1])], names=['Species', 'Family', 'Order', 'ZooMS_taxon'])).reset_index() all_d = pd.concat([all_d, d], axis=0) ranks_df = ranks.sort_values(by='Correlation', ascending=False) return all_d
但执行transform(df)时返回如下错误:
TypeError Traceback (most recent call last) <ipython-input-27-1fbc44e09cc7> in <cell line: 1>() ----> 1 transform(df) <ipython-input-26-775d7a15bb53> in transform(df) 7 print(type(pi)) 8 print(type(pi[1])) ----> 9 pi = pi[~np.isnan(pi)] 10 resampled = np.zeros_like(x) 11 for p in df_copy['mass']: TypeError: ufunc 'isnan' not supported for the input types, and the inputs could not be safely coerced to any supported types according to the casting rule ''safe''
我检查了pi的类型为<class 'numpy.ndarray'>,pi[0]的类型为<class 'float'>,但仍触发上述错误。我的问题是:为何pi数组中的数据未被识别为数值和NaN?为何被归类为float?我是否忽略了导致该错误的明显原因?
核心原因
你遇到的问题本质是**ref_peaks数据框中存在字符串类型的"NaN"或其他非数值文本**,而非真正的numpy/pandas识别的NaN值。虽然你看到pi[0]是float类型,但数组中可能混杂了字符串(比如CSV里的空值被读成了空字符串,或者"NA"/"nan"这类文本),导致整个数组的dtype是object,而非纯数值型。np.isnan无法处理object类型数组里的字符串元素,因此报错。
解决步骤
强制转换
ref_peaks为数值型:
在读取CSV后,用pd.to_numeric把所有列转成数值,同时把无法转换的非数值(比如文本型的"NaN")转为真正的NaN:ref_peaks = pd.read_csv('/content/drive/MyDrive/Archaeology 2024/Thesis Re-runs/Edited_ZooMS_Values.csv', sep=",", encoding='UTF-8', header=None, names=fields2) # 转换所有列为数值型,错误值转为NaN ref_peaks = ref_peaks.apply(pd.to_numeric, errors='coerce') ref_peaks = ref_peaks.iloc[1:]这一步能确保
ref_peaks里只有数值和真正的NaN,数组dtype会变成float64而非object。检查数据类型:
转换后可以用print(ref_peaks.dtypes)确认所有列都是float64,再用print(pi.dtype)检查pi数组的类型,应该是float64而非object。优化原代码中的数组转换:
原代码里pi = np.array(ref_peaks.iloc[i].to_numpy())可以简化为pi = ref_peaks.iloc[i].dropna().to_numpy(),直接跳过NaN值,省去np.isnan的判断步骤,更高效也避免错误。
额外注意点
- 读取CSV时,确保
fields2的列名数量和CSV的列数匹配,避免列错位导致的数据类型混乱。 - 检查
ref_taxa和ref_peaks合并后的references数据框,确保分类列是字符型,峰值列是数值型,避免后续索引错误。
内容的提问来源于stack exchange,提问作者Liz Quinlan

