如何在Pandas DataFrame列中应用自定义函数处理浮点数?
问题分析与解决
错误原因
你确实用错了apply():
df['location y'][index]取到的是单个numpy.float64数值,而apply()是Pandas的Series/DataFrame专属方法,单个数值没有这个属性,所以触发报错。- 完全没必要手动循环索引处理每一行,Pandas的核心设计就是批量处理整列/整表,循环反而会降低效率。
修正步骤
1. 先修复自定义函数的问题
你的函数存在两个小漏洞:
change_y里的[(num) - 470]是列表格式,会导致计算结果变成列表类型,要改成普通数值运算else分支没有返回值,会让未触发条件的数值变成NaN,必须返回原数值
修正后的函数:
def change_y(num): if num > 470: num = 470 - (num - 470) # 去掉方括号,改为普通减法 return num else: return num # 返回原数值,避免生成NaN def change_x(num): if num < 250: num = (250 - num) + 250 return num elif num > 250: num = 250 - (num - 250) return num else: return num # 返回原数值
2. 正确使用apply()处理整列
直接对整列调用apply(),无需循环索引:
# 先把字典转为真正的DataFrame(你原代码里的df是字典,需要先转换) import pandas as pd df = pd.DataFrame({ 'location x': [107.0, 254.0, 52.0, 640.0, 882.0], 'location y': [252.0, 56.0, 250.0, 86.0, 318.0] }) # 直接对整列应用函数,可选择覆盖原列或生成新列 df['location y'] = df['location y'].apply(change_y) df['location x'] = df['location x'].apply(change_x) # 查看处理结果 print(df)
3. 更高效的替代方案:矢量化操作
apply()本质还是逐行处理,针对大规模数据,用矢量化操作速度更快:
import numpy as np # 处理location x df['location x'] = np.where( df['location x'] < 250, (250 - df['location x']) + 250, np.where(df['location x'] > 250, 250 - (df['location x'] - 250), df['location x']) ) # 处理location y df['location y'] = np.where(df['location y'] > 470, 470 - (df['location y'] - 470), df['location y'])
最终处理结果
运行代码后得到的DataFrame:
location x location y 0 393.0 252.0 1 246.0 56.0 2 448.0 250.0 3 -140.0 86.0 4 -382.0 318.0
内容的提问来源于stack exchange,提问作者corey_h199
相关产品推荐
相关产品推荐

