You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用lambda+df.apply创建Pandas新列时遭遇内存错误求助

嘿,这个内存问题我太熟了!用df.apply(axis=1)处理24万行的数据集,确实容易因为逐行遍历的低效性导致内存爆掉。结合你的函数逻辑,给你几个实用的解决方案,从简单到进阶都有:

1. 用Pandas矢量化操作替代逐行apply(最推荐)

Pandas的矢量化操作是C级别的底层实现,比Python逐行循环快N倍,内存占用也小很多。我们可以把前两个字符串判断的分支用矢量化处理,只对需要处理列表的行用apply,大幅减少内存压力:

import numpy as np

# 先处理前两个字符串匹配的分支,生成布尔掩码
mask_string1 = df["column2"].str.contains("string1")
mask_string2 = df["column2"].str.contains("string2")

# 初始化新列,先处理前两种情况
df["new_col"] = np.where(mask_string1, df["column3"],
                        np.where(mask_string2, df["column1"], np.nan))

# 筛选出需要处理列表的行(也就是还没赋值的NaN行)
mask_need_list_process = df["new_col"].isna()

# 只对这些行调用你的列表处理逻辑,因为数量会少很多
def list_process_func(col1, col2, col3):
    # 这里写你处理列表的逻辑,比如遍历列表移除特定子串
    processed_list = [item.replace("target_substr", "") for item in col3]
    # 按你的实际逻辑返回对应值
    return processed_list  # 或者返回col1/col2的处理结果

df.loc[mask_need_list_process, "new_col"] = df.loc[mask_need_list_process].apply(
    lambda x: list_process_func(x["column1"], x["column2"], x["column3"]),
    axis=1
)

2. 用swifter库自动优化apply过程

如果你不想拆分函数逻辑,可以试试swifter库——它会自动判断你的函数是否能矢量化,不能的话就用Dask并行处理apply,还能优化内存占用:

首先安装库:

pip install swifter

然后替换你的apply代码:

import swifter

def func(col1, col2, col3):
    if "string1" in col2:
        return col3
    elif "string2" in col2:
        return col1
    else:
        # 这里写你的列表处理逻辑
        processed_list = [item.replace("target_substr", "") for item in col3]
        return processed_list

df["new_col"] = df.swifter.apply(
    lambda x: func(x["column1"], x["column2"], x["column3"]),
    axis=1
)

3. 分块处理超大数据集

如果你的数据集大到内存完全装不下,可以把DataFrame分成小块逐块处理,再合并结果:

# 设置分块大小,根据你的内存情况调整,比如1万行一块
chunk_size = 10000
processed_chunks = []

# 如果是从文件读取数据,直接用chunksize参数分块
for chunk in pd.read_csv("your_data_source.csv", chunksize=chunk_size):
    chunk["new_col"] = chunk.apply(
        lambda x: func(x["column1"], x["column2"], x["column3"]),
        axis=1
    )
    processed_chunks.append(chunk)

# 合并所有分块
df = pd.concat(processed_chunks, ignore_index=True)

4. 优化自定义函数的内存开销

针对你处理列表的分支,尽量用更高效的写法减少临时对象的创建:

  • 用列表推导式替代for循环+append,内存占用更小且更快
  • 避免在函数里创建不必要的大变量,尽量原地修改
  • 如果列表很长,考虑用生成器(不过Pandas列需要序列类型,所以可能还是列表更合适)

举个例子,把列表处理逻辑优化成:

def process_list(lst, target_substr):
    # 列表推导式直接生成处理后的列表,无需临时变量
    return [s.replace(target_substr, "") for s in lst if target_substr in s]

内容的提问来源于stack exchange,提问作者Lily

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 00:49:05