You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Colab文本逐行比较合并脚本性能优化求助

优化Google Colab平行语料处理脚本的速度与内存占用

核心性能瓶颈分析

原脚本运行缓慢、内存占用过高的原因集中在三点:

  • 线性查找的时间开销:每次判断english not in english_ukranian_en和调用index()都是O(n)级操作,循环执行m次(m为英波语料行数),整体时间复杂度达O(m*n),大语料场景下完全不可行。
  • 内存过载:一次性将所有语料文件加载到列表中,大文件会直接占满Colab内存,导致运行卡顿。
  • 频繁IO操作:每次匹配成功都重复打开/关闭输出文件,带来大量额外开销。

优化后的代码

下载解压部分(无需修改)

#@title Download and extract files.
!wget "https://object.pouta.csc.fi/OPUS-CCAligned/v1/moses/en-uk.txt.zip"
!wget "https://object.pouta.csc.fi/OPUS-CCAligned/v1/moses/en-pl.txt.zip"
!7z e "en-pl.txt.zip" -o"enpl"
!7z e "en-uk.txt.zip" -o"enuk"

核心处理逻辑(优化版)

#@title Do the task (Optimized Version)
from tqdm import tqdm

def build_en_to_target_map(en_path, target_path):
    """构建英文到目标语言的映射字典,O(n)时间完成"""
    en_to_target = {}
    # 逐行读取,避免一次性加载大文件到内存
    with open(en_path, "r", encoding="utf-8") as en_f, open(target_path, "r", encoding="utf-8") as target_f:
        for en_line, target_line in tqdm(zip(en_f, target_f), desc="Building mapping"):
            # 仅保留第一次出现的英文句子对应的目标语(与原脚本逻辑一致)
            if en_line not in en_to_target:
                en_to_target[en_line] = target_line
    return en_to_target

# 构建英文到乌克兰语的映射
en_to_uk = build_en_to_target_map("enuk/CCAligned.en-uk.en", "enuk/CCAligned.en-uk.uk")

# 提前打开输出文件,避免频繁IO操作
with open("pl-uk.pl", "w", encoding="utf-8") as pl_out, open("pl-uk.uk", "w", encoding="utf-8") as uk_out:
    # 逐行处理英波语料,内存占用极低
    with open("enpl/CCAligned.en-pl.en", "r", encoding="utf-8") as en_f, open("enpl/CCAligned.en-pl.pl", "r", encoding="utf-8") as pl_f:
        for en_line, pl_line in tqdm(zip(en_f, pl_f), desc="Processing en-pl pairs"):
            # O(1)时间快速查找匹配
            uk_line = en_to_uk.get(en_line)
            if uk_line is not None:
                pl_out.write(pl_line)
                uk_out.write(uk_line)

优化点说明

  1. 哈希表加速查找:用字典en_to_uk存储英文到乌克兰语的映射,将查找操作从O(n)降至O(1),整体时间复杂度优化为O(m + n),速度提升数个数量级。
  2. 逐行读取文件:不再一次性加载整个文件到内存,大幅降低内存占用,适配大语料处理场景。
  3. 减少IO开销:提前打开输出文件,全程仅执行一次打开/关闭操作,避免循环内的重复IO操作。
  4. 逻辑一致性:构建映射时仅保留英文句子第一次出现的对应乌克兰语,与原脚本index()取第一个匹配项的逻辑完全一致。

内容的提问来源于stack exchange,提问作者AgaZgaga

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 16:15:43