Pandas DataFrame应用自定义函数报错,如何解决类型错误?
问题分析与解决方案
嘿,我来帮你排查这个错误的原因,然后给出几种可行的解决办法!
错误根源
你遇到的TypeError: '>' not supported between instances of 'list' and 'int',问题出在**df.apply()的调用方式上**:
df['three'] = df.apply(func(['one'], ['two']), axis='columns')
这里你直接执行了func(['one'], ['two']),相当于把两个列表['one']和['two']作为参数传给了func,这时候func里的two变量是一个列表,自然没法和整数0做大小比较,所以报错了。
正确的apply用法是传入函数名,然后让Pandas自动把每一行的数据传递给函数,而不是提前把参数传给函数。
解决办法
方法1:修改自定义函数接收行对象
调整func,让它接收每一行的Series,直接提取对应列的值:
import pandas as pd df = pd.read_excel('test.xlsx') df.fillna(0.0, inplace=True) def func(row): three = row['one'] if row['two'] > 0: three = row['two'] return three # 直接传函数名,axis='columns'表示按行处理 df['three'] = df.apply(func, axis='columns')
方法2:用lambda在apply中传递列值
如果你不想修改原来的func,可以用lambda函数作为中间层,把每一行的one和two值传给func:
import pandas as pd df = pd.read_excel('test.xlsx') df.fillna(0.0, inplace=True) def func(one, two): three = one if two > 0: three = two return three # 用lambda提取每行的one和two,再传给func df['three'] = df.apply(lambda row: func(row['one'], row['two']), axis='columns')
方法3:矢量化操作(推荐)
对于这种简单的条件赋值,Pandas推荐用矢量化操作,比apply效率高很多(尤其是数据量大的时候),代码也更简洁:
import pandas as pd import numpy as np df = pd.read_excel('test.xlsx') df.fillna(0.0, inplace=True) # 用numpy.where实现条件赋值:满足条件取two,否则取one df['three'] = np.where(df['two'] > 0, df['two'], df['one']) # 或者用Pandas原生的where方法 # df['three'] = df['two'].where(df['two'] > 0, df['one'])
效果验证
执行上述任意一种方法后,你的DataFrame的three列会变成:
| one | two | three | |
|---|---|---|---|
| 0 | 12 | 0 | 12 |
| 1 | 34 | 0 | 34 |
| 2 | 56 | 34 | 34 |
| 3 | 67 | 5 | 5 |
| 4 | 45 | 0 | 45 |
内容的提问来源于stack exchange,提问作者wotesi
相关产品推荐
相关产品推荐

