Google Colab文本逐行比较合并脚本性能优化求助
优化Google Colab平行语料处理脚本的速度与内存占用
核心性能瓶颈分析
原脚本运行缓慢、内存占用过高的原因集中在三点:
- 线性查找的时间开销:每次判断
english not in english_ukranian_en和调用index()都是O(n)级操作,循环执行m次(m为英波语料行数),整体时间复杂度达O(m*n),大语料场景下完全不可行。 - 内存过载:一次性将所有语料文件加载到列表中,大文件会直接占满Colab内存,导致运行卡顿。
- 频繁IO操作:每次匹配成功都重复打开/关闭输出文件,带来大量额外开销。
优化后的代码
下载解压部分(无需修改)
#@title Download and extract files. !wget "https://object.pouta.csc.fi/OPUS-CCAligned/v1/moses/en-uk.txt.zip" !wget "https://object.pouta.csc.fi/OPUS-CCAligned/v1/moses/en-pl.txt.zip" !7z e "en-pl.txt.zip" -o"enpl" !7z e "en-uk.txt.zip" -o"enuk"
核心处理逻辑(优化版)
#@title Do the task (Optimized Version) from tqdm import tqdm def build_en_to_target_map(en_path, target_path): """构建英文到目标语言的映射字典,O(n)时间完成""" en_to_target = {} # 逐行读取,避免一次性加载大文件到内存 with open(en_path, "r", encoding="utf-8") as en_f, open(target_path, "r", encoding="utf-8") as target_f: for en_line, target_line in tqdm(zip(en_f, target_f), desc="Building mapping"): # 仅保留第一次出现的英文句子对应的目标语(与原脚本逻辑一致) if en_line not in en_to_target: en_to_target[en_line] = target_line return en_to_target # 构建英文到乌克兰语的映射 en_to_uk = build_en_to_target_map("enuk/CCAligned.en-uk.en", "enuk/CCAligned.en-uk.uk") # 提前打开输出文件,避免频繁IO操作 with open("pl-uk.pl", "w", encoding="utf-8") as pl_out, open("pl-uk.uk", "w", encoding="utf-8") as uk_out: # 逐行处理英波语料,内存占用极低 with open("enpl/CCAligned.en-pl.en", "r", encoding="utf-8") as en_f, open("enpl/CCAligned.en-pl.pl", "r", encoding="utf-8") as pl_f: for en_line, pl_line in tqdm(zip(en_f, pl_f), desc="Processing en-pl pairs"): # O(1)时间快速查找匹配 uk_line = en_to_uk.get(en_line) if uk_line is not None: pl_out.write(pl_line) uk_out.write(uk_line)
优化点说明
- 哈希表加速查找:用字典
en_to_uk存储英文到乌克兰语的映射,将查找操作从O(n)降至O(1),整体时间复杂度优化为O(m + n),速度提升数个数量级。 - 逐行读取文件:不再一次性加载整个文件到内存,大幅降低内存占用,适配大语料处理场景。
- 减少IO开销:提前打开输出文件,全程仅执行一次打开/关闭操作,避免循环内的重复IO操作。
- 逻辑一致性:构建映射时仅保留英文句子第一次出现的对应乌克兰语,与原脚本
index()取第一个匹配项的逻辑完全一致。
内容的提问来源于stack exchange,提问作者AgaZgaga
相关产品推荐
相关产品推荐

