You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的map函数优化DataFrame列表保存代码以提效

优化方案:用Map结合并行处理提速

首先要明确:单纯改用map函数本身不会缩短执行时间——你的原代码核心耗时是磁盘IO操作(创建文件夹、写入CSV),单线程下map和for循环的执行效率几乎一致。要真正提速,需要将map与多进程/多线程结合,并行处理IO任务。

步骤1:将循环逻辑封装为独立函数

先把原循环里的逻辑抽成可复用的函数,方便用map调用:

import os
import pandas as pd

def save_single_df(df):
    if df.empty:
        return
    
    # 构建目标路径(用os.path.join更规范,避免路径分隔符问题)
    base_dir = 'C:/Users/Desktop/'
    col1 = df.iloc[0]['Col1']
    col2 = df.iloc[0]['Col2']
    target_path = os.path.join(base_dir, col1, col2)
    
    # 创建文件夹(用exist_ok=True省去额外的路径判断,更高效)
    os.makedirs(target_path, exist_ok=True)
    
    # 处理DataFrame并保存
    day = df.iloc[0]['DAY']
    df_clean = df.drop(['Col1', 'Col2'], axis=1)
    csv_file = os.path.join(target_path, f'{day}.csv')
    df_clean.to_csv(csv_file, index=False)

步骤2:用普通Map替代For循环(无性能提升)

如果只是单纯替换循环为map,代码如下,但这和原for循环效率几乎一样:

# 单线程map,仅语法变化,无性能提升
list(map(save_single_df, Date))

步骤3:结合并行处理真正提速

因为磁盘IO是CPU等待型任务,可以用多进程/多线程并行执行,减少整体等待时间:

方案A:多进程(适合大数据量场景)

from multiprocessing import Pool

if __name__ == '__main__':
    # 进程数设为CPU核心数,避免调度开销过大
    with Pool(processes=os.cpu_count()) as pool:
        pool.map(save_single_df, Date)

方案B:多线程(IO密集型任务开销更小)

多线程的内存开销比多进程小,适合IO密集型场景:

from concurrent.futures import ThreadPoolExecutor

# 线程数可设为CPU核心数的2倍,充分利用IO等待时间
with ThreadPoolExecutor(max_workers=os.cpu_count() * 2) as executor:
    executor.map(save_single_df, Date)

额外优化点

  • 用os.path.join替代字符串拼接路径,避免不同系统的路径分隔符问题,也更健壮。
  • os.makedirs(target_path, exist_ok=True)直接替代if not os.path.exists: os.makedirs,减少一次磁盘IO判断,更高效。

内容的提问来源于stack exchange,提问作者Akansha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 17:40:28