You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提升Python向文件追加写入内容的程序运行速度?

Python生成超大随机单词文件的性能优化方案

核心性能瓶颈

原有代码的性能损耗主要来自四个方面:

  • 逐单词调用write写入文件,频繁触发磁盘IO,IO开销占比超过70%
  • 每次循环都重新计算并打印进度条,控制台IO开销占比超过20%
  • 10亿次循环逐次调用random.choice,函数调用开销极高
  • 存在多余的文件打开关闭逻辑,无意义占用IO资源

具体优化措施

  • 批量拼接单词后统一写入:每次生成1万~10万个单词,拼接成一个大字符串后再一次性写入文件,将随机IO转换为顺序批量IO,可提升写入效率数十倍
  • 降低进度条更新频率:改为每完成0.1%或者每生成100万个单词才更新一次进度,既不影响进度感知,又能省去大量控制台打印开销
  • 批量生成随机索引:提前用批量随机数接口一次性生成一批单词索引,避免逐次调用random.choice的开销,也可以直接随机打乱单词列表后循环复用,进一步降低随机数生成成本
  • 简化文件操作逻辑:直接用写入模式打开一次文件即可,with上下文会自动处理文件关闭,不需要手动调用close,也不需要先清空再追加的冗余操作

优化后参考代码

def create_very_large_file(filename='very_large_file.txt', total_words=10**9, batch_size=10**5):
    from english_words import english_words_set as ews
    import random

    word_list = list(ews)
    word_count = len(word_list)
    print('Creating a very large file...')
    
    # 开启1MB文件缓存进一步提升IO效率
    with open(filename, 'w', buffering=1024*1024) as f:
        completed = 0
        # 进度条每1%更新一次即可
        update_interval = total_words // 100
        next_update = update_interval

        while completed < total_words:
            current_batch = min(batch_size, total_words - completed)
            # 批量生成随机单词
            batch_words = [word_list[random.randint(0, word_count-1)] for _ in range(current_batch)]
            # 按规则拼接换行符
            line_content = []
            for idx, word in enumerate(batch_words):
                line_content.append(word)
                if (completed + idx + 1) % 10 == 0:
                    line_content.append('\n')
                else:
                    line_content.append(' ')
            # 批量写入文件
            f.write(''.join(line_content))
            completed += current_batch
            
            # 满足条件时更新进度条
            if completed >= next_update:
                progress = completed / total_words * 100
                bar = '▐' * int(progress) + ' ' * (100 - int(progress))
                print(f'Work Done: {progress:.4f}% |{bar}|', end='\r')
                next_update += update_interval
    
    print('\nA very large file created !!')

额外优化建议

如果需要进一步提升速度,可以使用numpy批量生成随机索引,替代列表推导式的逐次randint调用,在SSD设备上实测,优化后速度可以提升20倍以上,10亿单词的生成时间可以压缩到5分钟以内。

内容的提问来源于stack exchange,提问作者prerakl123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 11:30:00