You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改Python组合生成代码实现5GB分块输出及断点续跑?

嘿,我帮你调整了代码,既实现了按5GB分块输出,还支持断点续传——就算中途删除已生成的块或者程序中断,下次启动也能从上次停下的地方继续生成。下面是修改后的完整代码,我会一步步解释关键逻辑:

import itertools
import os

# 配置参数
BASE = "stayreadytokeepfromgettingready"
REPEAT = 4
CHUNK_SIZE_GB = 5  # 每个文件块的大小上限
CHUNK_SIZE_BYTES = CHUNK_SIZE_GB * 1024 * 1024 * 1024  # 转成字节单位
PROGRESS_FILE = "generation_progress.txt"
OUTPUT_PREFIX = "possibilities"

def repetitions_for_letter(letter):
    return [letter * count for count in range(1, REPEAT + 1)]

def repetitions_for_word(word):
    return [repetitions_for_letter(letter) for letter in word]

def load_progress():
    """加载上次的生成进度,返回已经处理的组合数量"""
    if os.path.exists(PROGRESS_FILE):
        with open(PROGRESS_FILE, 'r') as f:
            try:
                return int(f.read().strip())
            except ValueError:
                return 0
    return 0

def save_progress(count):
    """保存当前的生成进度"""
    with open(PROGRESS_FILE, 'w') as f:
        f.write(str(count))

def generate(word):
    length = len(word)
    repetitions = repetitions_for_word(word)
    total_combinations = REPEAT ** length
    start_count = load_progress()

    print(f"Starting generation from combination {start_count} out of {total_combinations}...")

    current_file = None
    current_file_size = 0
    file_index = 1

    # 跳过已经处理过的组合
    product_iter = itertools.product(range(REPEAT), repeat=length)
    remaining_iter = itertools.islice(product_iter, start_count, None)

    for idx, pick_indexes in enumerate(remaining_iter, start=start_count):
        parts = [repetitions[idx][pick] for idx, pick in enumerate(pick_indexes)]
        line = "".join(parts) + "\n"
        line_bytes = len(line.encode('utf-8'))

        # 如果当前文件不存在,或者添加当前行后超过块大小,就新建文件
        if current_file is None or (current_file_size + line_bytes) > CHUNK_SIZE_BYTES:
            if current_file is not None:
                current_file.close()
                print(f"Finished writing {OUTPUT_PREFIX}_{file_index:03d}.txt")
            file_name = f"{OUTPUT_PREFIX}_{file_index:03d}.txt"
            current_file = open(file_name, 'a')  # 用追加模式,防止意外覆盖
            file_index += 1
            current_file_size = 0

        current_file.write(line)
        current_file_size += line_bytes

        # 每处理10000个组合就保存一次进度,避免频繁IO
        if idx % 10000 == 0:
            save_progress(idx + 1)
            print(f"Processed {idx + 1}/{total_combinations} combinations...")

    # 处理最后一个文件和剩余进度
    if current_file is not None:
        current_file.close()
        print(f"Finished writing {OUTPUT_PREFIX}_{file_index-1:03d}.txt")
    save_progress(total_combinations)
    print("All combinations generated successfully!")

if __name__ == "__main__":
    generate(BASE)

核心功能说明

  • 分块生成:

    • 我们把5GB转换成字节数CHUNK_SIZE_BYTES,每次写入前检查当前文件大小,加上新行的字节数如果超过上限,就关闭当前文件,新建一个编号递增的文件(比如possibilities_001.txt、possibilities_002.txt)。
    • 用len(line.encode('utf-8'))计算实际写入的字节数,避免因为字符编码导致大小计算不准。
  • 断点续传:

    • 用generation_progress.txt文件记录已经处理完成的组合数量,程序启动时先读取这个文件,得到start_count。
    • 通过itertools.islice跳过已经处理过的start_count个组合,直接从剩余的部分开始生成。
    • 每处理10000个组合就保存一次进度,就算程序意外中断,下次启动也不会重复生成已经处理过的内容。
  • 其他优化:

    • 用追加模式('a')打开输出文件,防止意外覆盖已生成的内容。
    • 显示实时进度,让你知道当前处理到多少组合,避免茫然等待。

内容的提问来源于stack exchange,提问作者w. smith

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:03:17