Python多进程写入文件后文件未生成的问题咨询
Python多进程文件生成问题分析与解答
问题场景
编写了以下Python代码,意图通过多进程创建两个各包含5000行内容的文件,但执行后文件并未出现在工作目录中:
import multiprocessing import itertools import os # Define a generator function to yield lines of text def generate_lines(): # Replace this with your logic to generate lines of text for i in range(10000): yield f"Line {i + 1}" # Function to write lines to a file def write_lines(filename, lines): with open(filename, 'w') as file: try: for line in lines: file.write(line + '\n') except IOError as e: errno, strerror = e.args print(f"I/O error({errno}): {strerror}") if __name__ == '__main__': # Get the absolute path of the script's directory script_dir = os.path.dirname(os.path.abspath(__file__)) print('script_dir: ', script_dir) # Create a pool of processes` with multiprocessing.Pool(2) as pool: # Use itertools.islice to split the generator into chunks of 5000 lines each total_lines = 10000 chunk_size = 5000 lines_generator = generate_lines(total_lines) for i in range(0, total_lines // chunk_size): chunk = itertools.islice(lines_generator, i * chunk_size, (i + 1) * chunk_size) file_path = os.path.join(script_dir, f'file-{i}.txt') pool.apply_async(write_lines, (file_path, chunk)) # Wait for all processes to complete pool.close() pool.join() print("Writing completed successfully.")
一、文件未生成的底层原理
核心问题集中在三点:
- 函数参数不匹配:
generate_lines()定义时无参数,但主进程调用时传入了total_lines,会触发TypeError。但因为使用pool.apply_async提交异步任务,子进程的异常不会主动在主进程中打印,导致误以为代码正常执行。 - 生成器切片逻辑错误:即使参数问题修复,
itertools.islice对生成器的切片方式有误。第一次循环取前5000行后,生成器已被消费到第5000行位置;第二次循环用islice(lines_generator, 5000, 10000)时,起始位置是相对于当前生成器的偏移而非全局行数,实际取不到任何内容。 - 生成器无法跨进程序列化:在Windows系统下,
multiprocessing采用spawn方式创建子进程,需要将任务参数序列化后传递给子进程。但生成器(包括islice返回的迭代器)是带状态的对象,无法被pickle序列化,子进程会因此报错,无法执行写入操作。
二、文件是否被进程“吞掉”
文件并没有被进程“吞掉”,而是根本没有完成写入:
- 由于上述错误,子进程要么在启动时就因参数/序列化问题崩溃,要么无法获取有效行数据,导致文件要么未被创建,要么创建后为空(多数情况下因为异常提前退出,连文件都不会创建)。
- 主进程因为未捕获异步任务的异常,错误地打印了“Writing completed successfully.”,但实际子进程并未完成任务。
三、进程内存占用情况
- 代码中的错误会导致子进程启动后很快崩溃,因此内存占用极低,几乎可以忽略。
- 若修复错误,生成器本身是惰性计算的,每次仅生成一行数据,内存占用极低。在Unix系统(
fork方式创建进程)下,子进程会继承主进程的生成器状态,但依然保持低内存;Windows系统下需将生成器转为列表传递(会一次性占用存储10000行数据的内存,但总量很小),内存占用依然可控。
修复后的代码示例
import multiprocessing import os def generate_lines(total_lines): for i in range(total_lines): yield f"Line {i + 1}" def write_lines(filename, lines): with open(filename, 'w') as file: try: for line in lines: file.write(line + '\n') except IOError as e: errno, strerror = e.args print(f"I/O error({errno}): {strerror}") if __name__ == '__main__': script_dir = os.path.dirname(os.path.abspath(__file__)) print('script_dir: ', script_dir) total_lines = 10000 chunk_size = 5000 # 提前将生成器转为列表,确保可跨进程传递 all_lines = list(generate_lines(total_lines)) with multiprocessing.Pool(2) as pool: for i in range(0, total_lines // chunk_size): start = i * chunk_size end = start + chunk_size chunk = all_lines[start:end] file_path = os.path.join(script_dir, f'file-{i}.txt') pool.apply_async(write_lines, (file_path, chunk)) pool.close() pool.join() print("Writing completed successfully.")
内容的提问来源于stack exchange,提问作者Jwan622
相关产品推荐
相关产品推荐

