You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何减小PyInstaller/Nuitka生成的.exe文件体积?

缩减SRT翻译工具打包体积方案

问题背景

使用PyInstaller或Nuitka打包SRT翻译工具后,生成的.exe文件体积高达2.5GB。尝试调整打包配置、排除模块后体积有所减小,但会丢失必需依赖,计划通过用户端按需下载依赖的方式解决体积问题。

核心体积膨胀原因

打包体积过大的根源在于以下几部分被强制打包进exe:

  • PyTorch及其关联库(CUDA版本体积远超1GB)
  • Transformers库及相关依赖
  • facebook/nllb-200-3.3B预训练模型(本身体积接近10GB,若打包时误包含会大幅膨胀)

具体优化方案

1. 彻底剥离大型依赖,改为按需下载

打包时明确排除PyTorch、Transformers等体积较大的依赖,由用户首次运行工具时自动下载:

  • PyInstaller配置:
    pyinstaller --onefile --exclude-module torch --exclude-module transformers --exclude-module srt --strip --upx-dir=./upx your_script.py
    
  • Nuitka配置:
    nuitka --standalone --nofollow-import-to=torch --nofollow-import-to=transformers --nofollow-import-to=srt --strip your_script.py
    
  • 优化现有动态下载逻辑:
    • 将顶部的from transformers import AutoTokenizer, AutoModelForSeq2SeqLM移到get_translator函数内部,避免打包工具扫描到这些导入并尝试打包
    • 调整PyTorch下载命令,自动适配用户设备(检测CUDA版本,无CUDA则下载CPU版本):
      def download_pytorch():
          try:
              python_exe = sys.executable
              # 检测CUDA是否可用,自动选择对应PyTorch包
              cuda_check_cmd = [python_exe, "-c", "import torch; print(torch.cuda.is_available())"]
              cuda_available = subprocess.run(cuda_check_cmd, capture_output=True, text=True).stdout.strip() == "True"
              
              if cuda_available:
                  pip_command = [python_exe, "-m", "pip", "install", "--user", "torch", "torchvision", "torchaudio", "--extra-index-url", "https://download.pytorch.org/whl/cu118"]
              else:
                  pip_command = [python_exe, "-m", "pip", "install", "--user", "torch", "torchvision", "torchaudio"]
              
              result = subprocess.run(pip_command, stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True)
              # 后续逻辑保持不变
          except Exception as e:
              messagebox.showerror("Erro", f"Falha ao instalar PyTorch: {str(e)}")
              return False
      

2. 确保模型仅在用户端按需下载

当前代码已实现模型缓存到用户本地目录~/.cache/huggingface/hub,需确保打包时不会将模型文件或缓存目录包含进exe:

  • 打包时添加--exclude-dir排除Hugging Face缓存目录(PyInstaller):
    pyinstaller ... --exclude-dir="%USERPROFILE%\.cache\huggingface"
    
  • 在程序启动时增加模型检查逻辑,提前告知用户需要下载模型:
    def check_and_download_model(token):
        model_dir = os.path.join(os.path.expanduser("~"), ".cache", "huggingface", "hub", "models--facebook--nllb-200-3.3B")
        if not os.path.exists(model_dir):
            confirm = messagebox.askyesno("Model Not Found", "Translation model not detected. Do you want to download it now? (Size ~10GB)")
            if confirm:
                download_model_from_hub(token)
            else:
                messagebox.showerror("Error", "Translation cannot proceed without the model.")
                sys.exit(1)
    
    在process_translation函数开头调用该函数。

3. 打包后压缩与精简

  • 使用UPX压缩:对打包后的exe进行压缩,可进一步缩减30%-50%体积(部分PyTorch相关dll可能无法压缩,可单独排除)
  • 剥离调试信息:打包时添加--strip参数,移除二进制文件中的调试符号
  • 避免打包不必要的库:排除tkinter的冗余组件(比如tix等未使用的子模块),PyInstaller可通过--exclude-module tkinter.tix实现

代码关键调整点

  1. 将所有大型依赖的导入移到函数内部,避免打包工具扫描到:
    def get_translator(src_lang, tgt_lang):
        torch = import_pytorch_dynamically()
        if not torch:
            return None, None, None, None
        
        # 移到此处,避免打包时被扫描
        from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
        try:
            model_name = "facebook/nllb-200-3.3B"
            tokenizer = AutoTokenizer.from_pretrained(model_name)
            # 后续逻辑保持不变
    
  2. 增加启动时的轻量依赖检查:
    if __name__ == "__main__":
        # 启动时检查核心轻量依赖,不存在则自动安装
        try:
            import srt
            from huggingface_hub import HfApi
        except ImportError:
            python_exe = sys.executable
            subprocess.run([python_exe, "-m", "pip", "install", "--user", "srt", "huggingface-hub"])
            # 重新导入依赖
            import srt
            from huggingface_hub import HfApi
        
        root = tk.Tk()
        # 后续GUI初始化逻辑保持不变
    

内容的提问来源于stack exchange,提问作者MaloneFreak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 08:32:02