使用memory_profiler分析时生成器版文件拆分程序执行时间变长问题
我最近折腾了个Python小工具,用来把大文本文件拆成多个小文件,写了两个版本:一个用列表实现,另一个用生成器。为了搞清楚两者的性能差异,我用memory_profiler做了内存分析,结果挺出乎意料的——生成器版本内存效率拉满,但执行时间居然比列表版本还长了。下面给大家看看具体实现和测试数据:
1. 基于列表的实现
from memory_profiler import profile @profile() def main(): file_name = input("Enter the full path of file you want to split into smaller inputFiles: ") input_file = open(file_name).readlines() num_lines_orig = len(input_file) parts = int(input("Enter the number of parts you want to split in: ")) output_files = [(file_name + str(i)) for i in range(1, parts + 1)] st = 0 p = int(num_lines_orig / parts) ed = p for i in range(parts-1): with open(output_files[i], "w") as OF: OF.writelines(input_file[st:ed]) st = ed ed = st + p with open(output_files[-1], "w") as OF: OF.writelines(input_file[st:]) if __name__ == "__main__": main()
带memory_profiler的运行结果
$ time py36 Splitting\ text\ files_BAD_usingLists.py Enter the full path of file you want to split into smaller inputFiles: /apps/nttech/rbhanot/Downloads/test.txt Enter the number of parts you want to split in: 3 Filename: Splitting text files_BAD_usingLists.py Line # Mem usage Increment Line Contents =============================================== 6 47.8 MiB 0.0 MiB @profile() 7 def main(): 8 47.8 MiB 0.0 MiB file_name = input("Enter the full path of file you want to split into smaller inputFiles: ") 9 107.3 MiB 59.5 MiB input_file = open(file_name).readlines() 10 107.3 MiB 0.0 MiB num_lines_orig = len(input_file) 11 107.3 MiB 0.0 MiB parts = int(input("Enter the number of parts you want to split in: ")) 12 107.3 MiB 0.0 MiB output_files = [(file_name + str(i)) for i in range(1, parts + 1)] 13 107.3 MiB 0.0 MiB st = 0 14 107.3 MiB 0.0 MiB p = int(num_lines_orig / parts) 15 107.3 MiB 0.0 MiB ed = p 16 108.1 MiB 0.7 MiB for i in range(parts-1): 17 107.6 MiB -0.5 MiB with open(output_files[i], "w") as OF: 18 108.1 MiB 0.5 MiB OF.writelines(input_file[st:ed]) 19 108.1 MiB 0.0 MiB st = ed 20 108.1 MiB 0.0 MiB ed = st + p 21 22 108.1 MiB 0.0 MiB with open(output_files[-1], "w") as OF: 23 108.1 MiB 0.0 MiB OF.writelines(input_file[st:]) real 0m6.115s user 0m0.764s sys 0m0.052s
不带profiler的运行结果
$ time py36 Splitting\ text\ files_BAD_usingLists.py Enter the full path of file you want to split into smaller inputFiles: /apps/nttech/rbhanot/Downloads/test.txt Enter the number of parts you want to split in: 3 real 0m5.916s user 0m0.696s sys 0m0.080s
2. 基于生成器的实现
from memory_profiler import profile @profile() def main(): file_name = input("Enter the full path of file you want to split into smaller inputFiles: ") input_file = open(file_name) num_lines_orig = sum(1 for _ in input_file) input_file.seek(0) parts = int(input("Enter the number of parts you want to split in: ")) output_files = ((file_name + str(i)) for i in range(1, parts + 1)) st = 0 p = int(num_lines_orig / parts) ed = p for i in range(parts-1): file = next(output_files) with open(file, "w") as OF: for _ in range(st, ed): OF.writelines(input_file.readline()) st = ed ed = st + p if num_lines_orig - ed < p: ed = st + (num_lines_orig - ed) + p else: ed = st + p file = next(output_files) with open(file, "w") as OF: for _ in range(st, ed): OF.writelines(input_file.readline()) if __name__ == "__main__": main()
带memory_profiler的运行结果
$ time py36 -m memory_profiler Splitting\ text\ files_GOOD_usingGenerators.py Enter the full path of file you want to split into smaller inputFiles: /apps/nttech/rbhanot/Downloads/test.txt Enter the number of parts you want to split in: 3 Filename: Splitting text files_GOOD_usingGenerators.py Line # Mem usage Increment Line Contents =============================================== 4 47.988 MiB 0.000 MiB @profile() 5 def main(): 6 47.988 MiB 0.000 MiB file_name = input("Enter the full path of file you want to split into smaller inputFiles: ") 7 47.988 MiB 0.000 MiB input_file = open(file_name) 8 47.988 MiB 0.000 MiB num_lines_orig = sum(1 for _ in input_file) 9 47.988 MiB 0.000 MiB input_file.seek(0) 10 47.988 MiB 0.000 MiB parts = int(input("Enter the number of parts you want to split in: ")) 11 48.703 MiB 0.715 MiB output_files = ((file_name + str(i)) for i in range(1, parts + 1)) 12 47.988 MiB -0.715 MiB st = 0 13 47.988 MiB 0.000 MiB p = int(num_lines_orig / parts) 14 47.988 MiB 0.000 MiB ed = p 15 48.703 MiB 0.715 MiB for i in range(parts-1): 16 48.703 MiB 0.000 MiB file = next(output_files) 17 48.703 MiB 0.000 MiB with open(file, "w") as OF: 18 48.703 MiB 0.000 MiB for _ in range(st, ed): 19 48.703 MiB 0.000 MiB OF.writelines(input_file.readline()) 20 21 48.703 MiB 0.000 MiB st = ed 22 48.703 MiB 0.000 MiB ed = st + p 23 48.703 MiB 0.000 MiB if num_lines_orig - ed < p: 24 48.703 MiB 0.000 MiB ed = st + (num_lines_orig - ed) + p 25 48.703 MiB 0.000 MiB else: 26 48.703 MiB 0.000 MiB ed = st + p 27 48.703 MiB 0.000 MiB file = next(output_files) 28 48.703 MiB 0.000 MiB with open(file, "w") as OF: 29 48.703 MiB 0.000 MiB for _ in range(st, ed): 30 48.703 MiB 0.000 MiB OF.writelines(input_file.readline())
结果分析
内存表现对比
- 列表版本:调用
readlines()后内存直接飙升59.5MiB,因为它把整个文件的所有行一次性加载到内存列表中,后续操作全在内存完成。 - 生成器版本:内存占用几乎稳定在48MiB左右,全程没有批量加载文件内容,输出文件名也是用生成器表达式按需生成,完全不会占用额外内存存储所有文件名。
生成器版本耗时更长的原因
生成器版本虽然内存友好,但执行时间更长,主要有两个核心因素:
- 额外的文件遍历:为了统计总行数,生成器版本先完整遍历了一次文件(
sum(1 for _ in input_file)),而列表版本直接通过len(input_file)就能获取行数,多了一次IO操作开销。 - 逐行IO的累积开销:列表版本用
writelines()批量写入切片内容,系统可以一次性处理大量数据,IO效率极高;而生成器版本是逐行读取、逐行写入,每次readline()和writelines()都会触发一次系统调用,频繁的小IO操作会累积大量额外耗时。
如果想兼顾内存效率和执行速度,可以优化生成器版本:比如每次读取一批行(而非单行),或者用itertools.islice批量获取行,这样既能控制内存占用,又能减少IO操作次数。
内容的提问来源于stack exchange,提问作者Rohit
相关产品推荐
相关产品推荐

