You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

调用transform函数触发np.isnan TypeError:输入类型不支持

问题

我正尝试改编一篇近期发表论文中的代码用于自有数据集,对应的.ipynb文件可从原项目获取。我已按照作者描述的R预处理流程格式化数据,这是我首次使用Python,若表述不清或有低级错误请见谅。

最初的问题出在读取鱼类参考分类数据集(原代码针对哺乳动物),我通过读取两个独立CSV文件(一个存参考峰值,一个存分类信息)并合并,得到包含Order、Family、Common Name、Species及肽段位置列的reference数据框,前4列为字符型,后续列为数值或NaN。

相关代码如下:

df = pd.read_csv("/content/drive/MyDrive/Archaeology 2024/Thesis Re-runs/PP_Combo_1/Combo_1_C66.csv", sep=',')
ref_taxa = pd.read_csv('/content/drive/MyDrive/Archaeology 2024/Thesis Re-runs/Taxa.csv', sep=",", encoding='UTF-8', header=None, names=fields)
ref_peaks = pd.read_csv('/content/drive/MyDrive/Archaeology 2024/Thesis Re-runs/Edited_ZooMS_Values.csv', sep=",", encoding='UTF-8', header=None, names=fields2)
ref_taxa = ref_taxa.iloc[1:]
ref_peaks = ref_peaks.iloc[1:]
references = pd.concat([ref_taxa,ref_peaks], axis=1)

我要运行的transform函数代码如下:

def transform(df):
    df_copy = df.copy() # making a copy so I dont have to reload the dataframe everytime
    #then changes are made to df_copy   
    x = np.arange(500, 3500, 0.1)
    pi = np.array(ref_peaks.iloc[0].to_numpy())
    pi = pi[~np.isnan(pi)]      
    resampled = np.zeros_like(x)
    for p in df_copy['mass']:
        resampled[np.logical_and(p-x>-0.3,p-x<1.3)] = 1
        
    resampled_smooth = gaussian_filter1d(resampled, 100)
    plt.plot(x, resampled, label='sample', linewidth=1)
    plt.plot(x, resampled_smooth, label='sample', linewidth=1)
    
    n_spec = ref_peaks.shape[0]
    corrs = np.zeros(n_spec)
    for i in range(n_spec):
        y = np.zeros_like(x)
        pi = np.array(ref_peaks.iloc[i].to_numpy())
        pi = pi[~np.isnan(pi)]
        for p in pi:
            y[np.logical_and(p-x>-0.3,p-x<1.3)] = 1
            corrs[i] = np.corrcoef(y, resampled)[0,1]
        ranks = pd.DataFrame({'Order':reference['Order'][1:],'Family':reference['Family'][1:], 'Species':reference['Species'][1:], 'ZooMS_taxon':reference['ZooMS_taxon'][1:], 'Correlation':corrs})
    
    idx = ranks.sort_values(by='Correlation', ascending=False).index[:6] - 1
    # index [6] because the first 6 matches of species will be determined -- number depends on how many of the first n most likely species you want to look at
    
    all_d = pd.DataFrame(columns=ref_peaks.columns)
    for i in idx:
        y = np.zeros_like(x)
        pi = np.array(ref_peaks.iloc[i].to_numpy())
        pi = pi[~np.isnan(pi)]
        for p in pi:
            y[np.logical_and(p-x>-0.3,p-x<1.3)] = 1
        
        plt.figure()
        plt.plot(x, y * resampled)
        plt.title(reference['Species'][i+1]);

        v, ix, iy = np.intersect1d((np.round(pi, decimals=1)*10).astype(int), (x[y*resampled==1]*10).astype(int), return_indices=True)
 
        d = pd.DataFrame({k:p for k, p in zip(ref_peaks.columns[ix],v/10)}, 
                 index=pd.MultiIndex.from_tuples([(reference['Species'][i+1], reference['Family'][i+1], reference['Order'][i+1], reference['ZooMS_taxon'][i+1])],
                                                 names=['Species', 'Family', 'Order', 'ZooMS_taxon'])).reset_index()

        all_d = pd.concat([all_d, d], axis=0)
    
    ranks_df = ranks.sort_values(by='Correlation', ascending=False)

    
    return all_d

但执行transform(df)时返回如下错误:

TypeError                                 Traceback (most recent call last)
<ipython-input-27-1fbc44e09cc7> in <cell line: 1>()
----> 1 transform(df)

<ipython-input-26-775d7a15bb53> in transform(df)
      7     print(type(pi))
      8     print(type(pi[1]))
----> 9     pi = pi[~np.isnan(pi)]
     10     resampled = np.zeros_like(x)
     11     for p in df_copy['mass']:

TypeError: ufunc 'isnan' not supported for the input types, and the inputs could not be safely coerced to any supported types according to the casting rule ''safe''

我检查了pi的类型为<class 'numpy.ndarray'>,pi[0]的类型为<class 'float'>,但仍触发上述错误。我的问题是:为何pi数组中的数据未被识别为数值和NaN?为何被归类为float?我是否忽略了导致该错误的明显原因?


解决方案

核心原因

你遇到的问题本质是**ref_peaks数据框中存在字符串类型的"NaN"或其他非数值文本**,而非真正的numpy/pandas识别的NaN值。虽然你看到pi[0]是float类型,但数组中可能混杂了字符串(比如CSV里的空值被读成了空字符串,或者"NA"/"nan"这类文本),导致整个数组的dtype是object,而非纯数值型。np.isnan无法处理object类型数组里的字符串元素,因此报错。

解决步骤

  1. 强制转换ref_peaks为数值型:
    在读取CSV后,用pd.to_numeric把所有列转成数值,同时把无法转换的非数值(比如文本型的"NaN")转为真正的NaN:

    ref_peaks = pd.read_csv('/content/drive/MyDrive/Archaeology 2024/Thesis Re-runs/Edited_ZooMS_Values.csv', sep=",", encoding='UTF-8', header=None, names=fields2)
    # 转换所有列为数值型,错误值转为NaN
    ref_peaks = ref_peaks.apply(pd.to_numeric, errors='coerce')
    ref_peaks = ref_peaks.iloc[1:]
    

    这一步能确保ref_peaks里只有数值和真正的NaN,数组dtype会变成float64而非object。

  2. 检查数据类型:
    转换后可以用print(ref_peaks.dtypes)确认所有列都是float64,再用print(pi.dtype)检查pi数组的类型,应该是float64而非object。

  3. 优化原代码中的数组转换:
    原代码里pi = np.array(ref_peaks.iloc[i].to_numpy())可以简化为pi = ref_peaks.iloc[i].dropna().to_numpy(),直接跳过NaN值,省去np.isnan的判断步骤,更高效也避免错误。

额外注意点

  • 读取CSV时,确保fields2的列名数量和CSV的列数匹配,避免列错位导致的数据类型混乱。
  • 检查ref_taxa和ref_peaks合并后的references数据框,确保分类列是字符型,峰值列是数值型,避免后续索引错误。

内容的提问来源于stack exchange,提问作者Liz Quinlan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 06:34:57