使用lambda+df.apply创建Pandas新列时遭遇内存错误求助
嘿,这个内存问题我太熟了!用df.apply(axis=1)处理24万行的数据集,确实容易因为逐行遍历的低效性导致内存爆掉。结合你的函数逻辑,给你几个实用的解决方案,从简单到进阶都有:
1. 用Pandas矢量化操作替代逐行apply(最推荐)
Pandas的矢量化操作是C级别的底层实现,比Python逐行循环快N倍,内存占用也小很多。我们可以把前两个字符串判断的分支用矢量化处理,只对需要处理列表的行用apply,大幅减少内存压力:
import numpy as np # 先处理前两个字符串匹配的分支,生成布尔掩码 mask_string1 = df["column2"].str.contains("string1") mask_string2 = df["column2"].str.contains("string2") # 初始化新列,先处理前两种情况 df["new_col"] = np.where(mask_string1, df["column3"], np.where(mask_string2, df["column1"], np.nan)) # 筛选出需要处理列表的行(也就是还没赋值的NaN行) mask_need_list_process = df["new_col"].isna() # 只对这些行调用你的列表处理逻辑,因为数量会少很多 def list_process_func(col1, col2, col3): # 这里写你处理列表的逻辑,比如遍历列表移除特定子串 processed_list = [item.replace("target_substr", "") for item in col3] # 按你的实际逻辑返回对应值 return processed_list # 或者返回col1/col2的处理结果 df.loc[mask_need_list_process, "new_col"] = df.loc[mask_need_list_process].apply( lambda x: list_process_func(x["column1"], x["column2"], x["column3"]), axis=1 )
2. 用swifter库自动优化apply过程
如果你不想拆分函数逻辑,可以试试swifter库——它会自动判断你的函数是否能矢量化,不能的话就用Dask并行处理apply,还能优化内存占用:
首先安装库:
pip install swifter
然后替换你的apply代码:
import swifter def func(col1, col2, col3): if "string1" in col2: return col3 elif "string2" in col2: return col1 else: # 这里写你的列表处理逻辑 processed_list = [item.replace("target_substr", "") for item in col3] return processed_list df["new_col"] = df.swifter.apply( lambda x: func(x["column1"], x["column2"], x["column3"]), axis=1 )
3. 分块处理超大数据集
如果你的数据集大到内存完全装不下,可以把DataFrame分成小块逐块处理,再合并结果:
# 设置分块大小,根据你的内存情况调整,比如1万行一块 chunk_size = 10000 processed_chunks = [] # 如果是从文件读取数据,直接用chunksize参数分块 for chunk in pd.read_csv("your_data_source.csv", chunksize=chunk_size): chunk["new_col"] = chunk.apply( lambda x: func(x["column1"], x["column2"], x["column3"]), axis=1 ) processed_chunks.append(chunk) # 合并所有分块 df = pd.concat(processed_chunks, ignore_index=True)
4. 优化自定义函数的内存开销
针对你处理列表的分支,尽量用更高效的写法减少临时对象的创建:
- 用列表推导式替代for循环+append,内存占用更小且更快
- 避免在函数里创建不必要的大变量,尽量原地修改
- 如果列表很长,考虑用生成器(不过Pandas列需要序列类型,所以可能还是列表更合适)
举个例子,把列表处理逻辑优化成:
def process_list(lst, target_substr): # 列表推导式直接生成处理后的列表,无需临时变量 return [s.replace(target_substr, "") for s in lst if target_substr in s]
内容的提问来源于stack exchange,提问作者Lily
相关产品推荐
相关产品推荐

