You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas使用Lambda计算剩余租期遇内存错误,寻求替代实现方法

问题描述

我有一个包含多列的Pandas DataFrame,其中remaining_lease列存在75%的NaN值,不想删除该列,希望通过lease_commence_date和current_year两列计算来填充该列的NaN值,计算公式为:
remaining_lease = 99 - ( current_year - lease_commence_date)
示例:当current_year = 2022且lease_commence_date = 1979时,remaining_lease = 99 - (2022 - 1979) = 56

我编写了如下函数实现该逻辑:

import math
def remaining_lease_year(x, current_year, commense_year):
    if math.isnan(x): # if the value is nan
        lease_year = 99 - (current_year - commense_year)
        return lease_year
    else: #if the value is not nan
        return x

df['remaining_lease'] = df['remaining_lease'].apply(lambda x: remaining_lease_year(x, df['current_year'], df['lease_commence_date']))

运行时出现错误:

MemoryError: Unable to allocate 7.08 MiB for an array with shape (927465,) and data type int64

请问是否有其他可行的实现方式?

解决方案

问题根源是你在apply的lambda里传入了整个current_year和lease_commence_date Series,而非对应行的单个值,导致每次函数调用都要处理全量数组,触发内存溢出。推荐使用Pandas原生的向量化操作,效率更高且内存友好:

方法1:使用fillna结合向量化计算

先计算出所有需要填充的值,再批量替换NaN:

# 生成填充用的数值列
fill_values = 99 - (df['current_year'] - df['lease_commence_date'])
# 填充remaining_lease列的NaN值
df['remaining_lease'] = df['remaining_lease'].fillna(fill_values)

方法2:通过loc定位NaN行批量赋值

直接筛选出remaining_lease为NaN的行,对这些行批量计算并赋值:

# 生成NaN值的掩码
nan_mask = df['remaining_lease'].isna()
# 对掩码选中的行赋值
df.loc[nan_mask, 'remaining_lease'] = 99 - (df.loc[nan_mask, 'current_year'] - df.loc[nan_mask, 'lease_commence_date'])

这两种方式都是基于Pandas底层的向量化运算,不需要逐行调用Python函数,既避免了内存问题,计算速度也远快于apply。

内容的提问来源于stack exchange,提问作者Vishwas Basotra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 14:56:18