You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用memory_profiler分析时生成器版文件拆分程序执行时间变长问题

我最近折腾了个Python小工具,用来把大文本文件拆成多个小文件,写了两个版本:一个用列表实现,另一个用生成器。为了搞清楚两者的性能差异,我用memory_profiler做了内存分析,结果挺出乎意料的——生成器版本内存效率拉满,但执行时间居然比列表版本还长了。下面给大家看看具体实现和测试数据:


1. 基于列表的实现

from memory_profiler import profile
@profile()
def main():
    file_name = input("Enter the full path of file you want to split into smaller inputFiles: ")
    input_file = open(file_name).readlines()
    num_lines_orig = len(input_file)
    parts = int(input("Enter the number of parts you want to split in: "))
    output_files = [(file_name + str(i)) for i in range(1, parts + 1)]
    st = 0
    p = int(num_lines_orig / parts)
    ed = p
    for i in range(parts-1):
        with open(output_files[i], "w") as OF:
            OF.writelines(input_file[st:ed])
        st = ed
        ed = st + p
    with open(output_files[-1], "w") as OF:
        OF.writelines(input_file[st:])
if __name__ == "__main__":
    main()

带memory_profiler的运行结果

$ time py36 Splitting\ text\ files_BAD_usingLists.py
Enter the full path of file you want to split into smaller inputFiles: /apps/nttech/rbhanot/Downloads/test.txt
Enter the number of parts you want to split in: 3
Filename: Splitting text files_BAD_usingLists.py
Line #    Mem usage    Increment  Line Contents
===============================================
     6     47.8 MiB     0.0 MiB  @profile()
     7                         def main():
     8     47.8 MiB     0.0 MiB      file_name = input("Enter the full path of file you want to split into smaller inputFiles: ")
     9    107.3 MiB    59.5 MiB      input_file = open(file_name).readlines()
    10    107.3 MiB     0.0 MiB      num_lines_orig = len(input_file)
    11    107.3 MiB     0.0 MiB      parts = int(input("Enter the number of parts you want to split in: "))
    12    107.3 MiB     0.0 MiB      output_files = [(file_name + str(i)) for i in range(1, parts + 1)]
    13    107.3 MiB     0.0 MiB      st = 0
    14    107.3 MiB     0.0 MiB      p = int(num_lines_orig / parts)
    15    107.3 MiB     0.0 MiB      ed = p
    16    108.1 MiB     0.7 MiB      for i in range(parts-1):
    17    107.6 MiB    -0.5 MiB          with open(output_files[i], "w") as OF:
    18    108.1 MiB     0.5 MiB              OF.writelines(input_file[st:ed])
    19    108.1 MiB     0.0 MiB          st = ed
    20    108.1 MiB     0.0 MiB          ed = st + p
    21
    22    108.1 MiB     0.0 MiB      with open(output_files[-1], "w") as OF:
    23    108.1 MiB     0.0 MiB          OF.writelines(input_file[st:])

real    0m6.115s
user    0m0.764s
sys     0m0.052s

不带profiler的运行结果

$ time py36 Splitting\ text\ files_BAD_usingLists.py
Enter the full path of file you want to split into smaller inputFiles: /apps/nttech/rbhanot/Downloads/test.txt
Enter the number of parts you want to split in: 3

real    0m5.916s
user    0m0.696s
sys     0m0.080s

2. 基于生成器的实现

from memory_profiler import profile
@profile()
def main():
    file_name = input("Enter the full path of file you want to split into smaller inputFiles: ")
    input_file = open(file_name)
    num_lines_orig = sum(1 for _ in input_file)
    input_file.seek(0)
    parts = int(input("Enter the number of parts you want to split in: "))
    output_files = ((file_name + str(i)) for i in range(1, parts + 1))
    st = 0
    p = int(num_lines_orig / parts)
    ed = p
    for i in range(parts-1):
        file = next(output_files)
        with open(file, "w") as OF:
            for _ in range(st, ed):
                OF.writelines(input_file.readline())
        st = ed
        ed = st + p
    if num_lines_orig - ed < p:
        ed = st + (num_lines_orig - ed) + p
    else:
        ed = st + p
    file = next(output_files)
    with open(file, "w") as OF:
        for _ in range(st, ed):
            OF.writelines(input_file.readline())
if __name__ == "__main__":
    main()

带memory_profiler的运行结果

$ time py36 -m memory_profiler Splitting\ text\ files_GOOD_usingGenerators.py
Enter the full path of file you want to split into smaller inputFiles: /apps/nttech/rbhanot/Downloads/test.txt
Enter the number of parts you want to split in: 3
Filename: Splitting text files_GOOD_usingGenerators.py
Line #    Mem usage    Increment  Line Contents
===============================================
     4     47.988 MiB     0.000 MiB  @profile()
     5                         def main():
     6     47.988 MiB     0.000 MiB      file_name = input("Enter the full path of file you want to split into smaller inputFiles: ")
     7     47.988 MiB     0.000 MiB      input_file = open(file_name)
     8     47.988 MiB     0.000 MiB      num_lines_orig = sum(1 for _ in input_file)
     9     47.988 MiB     0.000 MiB      input_file.seek(0)
    10     47.988 MiB     0.000 MiB      parts = int(input("Enter the number of parts you want to split in: "))
    11     48.703 MiB     0.715 MiB      output_files = ((file_name + str(i)) for i in range(1, parts + 1))
    12     47.988 MiB    -0.715 MiB      st = 0
    13     47.988 MiB     0.000 MiB      p = int(num_lines_orig / parts)
    14     47.988 MiB     0.000 MiB      ed = p
    15     48.703 MiB     0.715 MiB      for i in range(parts-1):
    16     48.703 MiB     0.000 MiB          file = next(output_files)
    17     48.703 MiB     0.000 MiB          with open(file, "w") as OF:
    18     48.703 MiB     0.000 MiB              for _ in range(st, ed):
    19     48.703 MiB     0.000 MiB                  OF.writelines(input_file.readline())
    20
    21     48.703 MiB     0.000 MiB      st = ed
    22     48.703 MiB     0.000 MiB      ed = st + p
    23     48.703 MiB     0.000 MiB      if num_lines_orig - ed < p:
    24     48.703 MiB     0.000 MiB          ed = st + (num_lines_orig - ed) + p
    25     48.703 MiB     0.000 MiB      else:
    26     48.703 MiB     0.000 MiB          ed = st + p
    27     48.703 MiB     0.000 MiB      file = next(output_files)
    28     48.703 MiB     0.000 MiB      with open(file, "w") as OF:
    29     48.703 MiB     0.000 MiB          for _ in range(st, ed):
    30     48.703 MiB     0.000 MiB              OF.writelines(input_file.readline())

结果分析

内存表现对比

  • 列表版本:调用readlines()后内存直接飙升59.5MiB,因为它把整个文件的所有行一次性加载到内存列表中,后续操作全在内存完成。
  • 生成器版本:内存占用几乎稳定在48MiB左右,全程没有批量加载文件内容,输出文件名也是用生成器表达式按需生成,完全不会占用额外内存存储所有文件名。

生成器版本耗时更长的原因

生成器版本虽然内存友好,但执行时间更长,主要有两个核心因素:

  1. 额外的文件遍历:为了统计总行数,生成器版本先完整遍历了一次文件(sum(1 for _ in input_file)),而列表版本直接通过len(input_file)就能获取行数,多了一次IO操作开销。
  2. 逐行IO的累积开销:列表版本用writelines()批量写入切片内容,系统可以一次性处理大量数据,IO效率极高;而生成器版本是逐行读取、逐行写入,每次readline()和writelines()都会触发一次系统调用,频繁的小IO操作会累积大量额外耗时。

如果想兼顾内存效率和执行速度,可以优化生成器版本:比如每次读取一批行(而非单行),或者用itertools.islice批量获取行,这样既能控制内存占用,又能减少IO操作次数。

内容的提问来源于stack exchange,提问作者Rohit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 09:00:23