You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多进程写入文件后文件未生成的问题咨询

Python多进程文件生成问题分析与解答

问题场景

编写了以下Python代码,意图通过多进程创建两个各包含5000行内容的文件,但执行后文件并未出现在工作目录中:

import multiprocessing
import itertools
import os


# Define a generator function to yield lines of text
def generate_lines():
    # Replace this with your logic to generate lines of text
    for i in range(10000):
        yield f"Line {i + 1}"


# Function to write lines to a file
def write_lines(filename, lines):
    with open(filename, 'w') as file:
        try:
            for line in lines:
                file.write(line + '\n')
        except IOError as e:
            errno, strerror = e.args
            print(f"I/O error({errno}): {strerror}")


if __name__ == '__main__':
    # Get the absolute path of the script's directory
    script_dir = os.path.dirname(os.path.abspath(__file__))
    print('script_dir: ', script_dir)
    # Create a pool of processes`
    with multiprocessing.Pool(2) as pool:
        # Use itertools.islice to split the generator into chunks of 5000 lines each
        total_lines = 10000
        chunk_size = 5000
        lines_generator = generate_lines(total_lines)
        for i in range(0, total_lines // chunk_size):
            chunk = itertools.islice(lines_generator, i * chunk_size, (i + 1) * chunk_size)
            file_path = os.path.join(script_dir, f'file-{i}.txt')
            pool.apply_async(write_lines, (file_path, chunk))

        # Wait for all processes to complete
        pool.close()
        pool.join()

    print("Writing completed successfully.")

一、文件未生成的底层原理

核心问题集中在三点:

  1. 函数参数不匹配:generate_lines()定义时无参数,但主进程调用时传入了total_lines,会触发TypeError。但因为使用pool.apply_async提交异步任务,子进程的异常不会主动在主进程中打印,导致误以为代码正常执行。
  2. 生成器切片逻辑错误:即使参数问题修复,itertools.islice对生成器的切片方式有误。第一次循环取前5000行后,生成器已被消费到第5000行位置;第二次循环用islice(lines_generator, 5000, 10000)时,起始位置是相对于当前生成器的偏移而非全局行数,实际取不到任何内容。
  3. 生成器无法跨进程序列化:在Windows系统下,multiprocessing采用spawn方式创建子进程,需要将任务参数序列化后传递给子进程。但生成器(包括islice返回的迭代器)是带状态的对象,无法被pickle序列化,子进程会因此报错,无法执行写入操作。

二、文件是否被进程“吞掉”

文件并没有被进程“吞掉”,而是根本没有完成写入:

  • 由于上述错误,子进程要么在启动时就因参数/序列化问题崩溃,要么无法获取有效行数据,导致文件要么未被创建,要么创建后为空(多数情况下因为异常提前退出,连文件都不会创建)。
  • 主进程因为未捕获异步任务的异常,错误地打印了“Writing completed successfully.”,但实际子进程并未完成任务。

三、进程内存占用情况

  • 代码中的错误会导致子进程启动后很快崩溃,因此内存占用极低,几乎可以忽略。
  • 若修复错误,生成器本身是惰性计算的,每次仅生成一行数据,内存占用极低。在Unix系统(fork方式创建进程)下,子进程会继承主进程的生成器状态,但依然保持低内存;Windows系统下需将生成器转为列表传递(会一次性占用存储10000行数据的内存,但总量很小),内存占用依然可控。

修复后的代码示例

import multiprocessing
import os


def generate_lines(total_lines):
    for i in range(total_lines):
        yield f"Line {i + 1}"


def write_lines(filename, lines):
    with open(filename, 'w') as file:
        try:
            for line in lines:
                file.write(line + '\n')
        except IOError as e:
            errno, strerror = e.args
            print(f"I/O error({errno}): {strerror}")


if __name__ == '__main__':
    script_dir = os.path.dirname(os.path.abspath(__file__))
    print('script_dir: ', script_dir)
    
    total_lines = 10000
    chunk_size = 5000
    # 提前将生成器转为列表,确保可跨进程传递
    all_lines = list(generate_lines(total_lines))
    
    with multiprocessing.Pool(2) as pool:
        for i in range(0, total_lines // chunk_size):
            start = i * chunk_size
            end = start + chunk_size
            chunk = all_lines[start:end]
            file_path = os.path.join(script_dir, f'file-{i}.txt')
            pool.apply_async(write_lines, (file_path, chunk))
        
        pool.close()
        pool.join()

    print("Writing completed successfully.")

内容的提问来源于stack exchange,提问作者Jwan622

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 21:47:02