You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas DataFrame处理后返回NoneType,无法访问问题求助

Pandas数据处理函数返回None的问题与修复

问题描述

从PDF提取表格生成Pandas DataFrame后,编写了4个互相调用的处理函数,但调用顶层函数data_cleaner后返回NoneType,无法获取处理后的结果。尽管在最后一个函数中能看到完全更新的DataFrame,但返回后无法访问。

用户原始代码:

def data_cleaner(dataFrame):
    #removing random rows
    removed = dataFrame.drop(columns=['Unnamed: 1','Unnamed: 2','Unnamed: 4','Unnamed: 5','Unnamed: 7','Unnamed: 9','Unnamed: 11','Unnamed: 13','Unnamed: 15','Unnamed: 17','Unnamed: 19'])
    #call next method
    col_combiner(removed)

def col_combiner(dataFrame):
    
    #Grabbing first and second row of table to combine
    first_row = dataFrame.iloc[0]
    second_row = dataFrame.iloc[1]
    #List to combine columns
    newColNames = []
    #Run through each row and combine them into one name
    for i,j in zip(first_row,second_row):
        #Check to see if they are not strings, if they are not convert it
        if not isinstance(i,str):
            i = str(i)
        if not isinstance(j,str):
            j = str(j)
        newString = ''
        #Check for double NAN case and change it to Expenses
        if i == 'nan' and j == 'nan':
            i = 'Expenses'
            newString = newString + i
        #Check for leading NAN and remove it
        elif i == 'nan':
            newString = newString + j
        else:            
            newString = newString + i + ' ' + j
            
    
        newColNames.append(newString)
    
    #Now update the dataframes column names
    dataFrame.columns = newColNames
    
    #Remove the name rows since they are now the column names
    dataFrame = dataFrame.iloc[2:,:]
    
    #Going to clean the values in the DF
    clean_numbers(dataFrame)


def clean_numbers(dataFrame):
    #Fill NAN values with 0
    noNan = dataFrame.fillna(0)
    
    #Pull each column, clean the values, then put it back
    for i in range(noNan.shape[1]):
        colList = noNan.iloc[:,i].tolist()
        #calling to clean the column so that it is all ints
        col_checker(colList)
        noNan.iloc[:,i] = colList
    
    
    return noNan

def col_checker(col):
    #Going through, checking and cleaning
    for i in range(len(col)):
        #print(type(colList[i]))
        if isinstance(col[i],str):
            col[i] = col[i].replace(',','')
            if col[i].isdigit():
                #print('not here')
                col[i] = int(col[i]) 
            #If it is not a number then make it 0
            else:
                col[i] = 0

调用代码:

doesThisWork = data_cleaner(cleaner)
type(doesThisWork)  # 返回 NoneType

核心问题分析

所有问题的根源是顶层和中间函数没有正确返回处理后的结果:

  • data_cleaner调用col_combiner后没有返回其结果
  • col_combiner调用clean_numbers后没有返回其结果
  • 另外,col_combiner中dataFrame = dataFrame.iloc[2:,:]生成了新的DataFrame对象,必须把这个新对象传递给clean_numbers

修复后的代码

def data_cleaner(dataFrame):
    # 移除指定列
    removed = dataFrame.drop(columns=['Unnamed: 1','Unnamed: 2','Unnamed: 4','Unnamed: 5','Unnamed: 7','Unnamed: 9','Unnamed: 11','Unnamed: 13','Unnamed: 15','Unnamed: 17','Unnamed: 19'])
    # 返回col_combiner的处理结果
    return col_combiner(removed)

def col_combiner(dataFrame):
    first_row = dataFrame.iloc[0]
    second_row = dataFrame.iloc[1]
    newColNames = []
    
    for i,j in zip(first_row,second_row):
        i = str(i) if not isinstance(i, str) else i
        j = str(j) if not isinstance(j, str) else j
        
        if i == 'nan' and j == 'nan':
            newString = 'Expenses'
        elif i == 'nan':
            newString = j
        else:            
            newString = f"{i} {j}"
            
        newColNames.append(newString)
    
    # 更新列名
    dataFrame.columns = newColNames
    # 移除前两行,生成新的DataFrame
    trimmed_df = dataFrame.iloc[2:,:]
    # 返回clean_numbers的处理结果
    return clean_numbers(trimmed_df)

def clean_numbers(dataFrame):
    noNan = dataFrame.fillna(0)
    
    for i in range(noNan.shape[1]):
        colList = noNan.iloc[:,i].tolist()
        col_checker(colList)
        noNan.iloc[:,i] = colList
    
    return noNan

def col_checker(col):
    for idx in range(len(col)):
        if isinstance(col[idx], str):
            cleaned_val = col[idx].replace(',', '')
            col[idx] = int(cleaned_val) if cleaned_val.isdigit() else 0

关键修改点

  1. data_cleaner函数:添加return col_combiner(removed),将下层函数的结果向上传递
  2. col_combiner函数:
    • 将dataFrame.iloc[2:,:]赋值给新变量trimmed_df,避免混淆原对象
    • 添加return clean_numbers(trimmed_df),传递处理后的新DataFrame并返回结果
    • 简化字符串拼接逻辑,使用f-string让代码更简洁
  3. col_checker函数:提取中间变量cleaned_val,减少重复操作,提升可读性

额外优化建议(针对Python新手)

  • 尽量避免在Pandas中用循环处理列,可使用apply或向量化操作提升效率,比如clean_numbers可以改写为:
    def clean_numbers(dataFrame):
        def clean_val(val):
            if isinstance(val, str):
                val = val.replace(',', '')
                return int(val) if val.isdigit() else 0
            return val if pd.notna(val) else 0
        
        return dataFrame.fillna(0).applymap(clean_val)
    
  • 检查NaN值时,建议使用Pandas内置的pd.isna()而非字符串比较i == 'nan',因为实际的NaN是float类型,字符串比较可能出现误判:
    # 在col_combiner中替换原判断逻辑
    if pd.isna(i) and pd.isna(j):
        newString = 'Expenses'
    elif pd.isna(i):
        newString = str(j)
    else:
        newString = f"{i} {j}"
    

内容的提问来源于stack exchange,提问作者db1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 18:33:34