You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python处理Minecraft模组大文件转二进制时触发Memory Error

解决Minecraft整合包大文件二进制转储时的MemoryError问题

我正在开发一个程序,用来为Minecraft整合包生成包含各类文件二进制数据的字典。小文件处理正常,但遇到90MB的“Better End Reforked”这类大模组文件时,程序直接报MemoryError崩溃。

我的代码:

from os      import listdir
from os.path import abspath, splitext, basename, getsize, isfile, join, expandvars

def getBool(statement):
    return (True if statement else False)

def getFiles(path):
    list = []

    for file in listdir(path):
        filePath = join(path, file)
        if isfile(filePath):
            list.append((file, abspath(filePath)))
    
    return list

def getIdentity(path):
    name = basename(path)
    
    return {
        'name': name,
        'basename': splitext(name)[0], 
        'extension': splitext(name)[1],
        'realpath': abspath(expandvars(path)),
        'size': getsize(path)
    }

def fromFileToBinary(path):
    try:
        with open(path, 'rb') as file:
            binary = file.read()
            return binary.hex()
    except IOError:
        return None


output = getIdentity('./output/bin.py')['realpath']

config       = 'C:/Users/Berdy Alexei/Downloads/modpack/optional/config'
mod          = 'C:/Users/Berdy Alexei/Downloads/modpack/optional/mods'
resourcepack = 'C:/Users/Berdy Alexei/Downloads/modpack/optional/resourcepacks'
script       = 'C:/Users/Berdy Alexei/Downloads/modpack/optional/scripts'

bool = getBool(config or mod or resourcepack or script)

def _bin(path = None, bool = True):
    dictionary = {}

    if path and bool:
        files = getFiles(path)
        
        for file in files:
            name, path = file
            dictionary[name] = fromFileToBinary(path)

    return dictionary

content = {
    'default': {
        'config':       _bin('C:/Users/Berdy Alexei/Downloads/modpack/default/config', False),
        'mods':         _bin('C:/Users/Berdy Alexei/Downloads/modpack/default/mods', False),
        'resourcepack': _bin('C:/Users/Berdy Alexei/Downloads/modpack/default/resourcepacks', False),
        'script':       _bin('C:/Users/Berdy Alexei/Downloads/modpack/default/scripts', False)
    },
    'optional': (bool, {
        'config':       _bin(config),
        'mods':         _bin(mod),
        'resourcepack': _bin(resourcepack),
        'script':       _bin(script)
    })
}

with open(output, "w") as file:
    file.write('BIN = {}'.format(content))

报错信息:

Traceback (most recent call last):
  File "c:\Folders\Archivos\Proyectos\InstallerCrafter\InstallerCrafter.py", line 75, in <module>
    file.write('BIN = {}'.format(content))
MemoryError

处理的文件:
文件列表截图


问题根源

  1. 一次性读取大文件:fromFileToBinary里用file.read()直接把90MB的文件全读进内存,转成hex字符串后体积会翻倍(每个字节转成两个字符),光是这一个文件就占180MB内存。
  2. 全量存储再写入:所有文件的hex字符串都存在content字典里,最后还要把整个字典转成一个巨大的字符串写入,内存直接被撑爆。

修复方案

1. 分块读取大文件转hex

修改fromFileToBinary,分块读取文件,避免一次性加载整个文件:

def fromFileToBinary(path):
    try:
        hex_str = []
        with open(path, 'rb') as file:
            # 每次读4KB块,平衡效率和内存占用
            chunk = file.read(4096)
            while chunk:
                hex_str.append(chunk.hex())
                chunk = file.read(4096)
        return ''.join(hex_str)
    except IOError:
        return None

2. 边处理边写入文件,避免全量存内存

不要先把所有数据塞进字典再写入,而是直接生成Python代码并逐段写入文件,这样内存只会保留当前处理的文件数据:

output = getIdentity('./output/bin.py')['realpath']

# 定义要处理的目录
dirs = {
    'default': {
        'config': ('C:/Users/Berdy Alexei/Downloads/modpack/default/config', False),
        'mods': ('C:/Users/Berdy Alexei/Downloads/modpack/default/mods', False),
        'resourcepack': ('C:/Users/Berdy Alexei/Downloads/modpack/default/resourcepacks', False),
        'script': ('C:/Users/Berdy Alexei/Downloads/modpack/default/scripts', False)
    },
    'optional': {
        'enabled': any([config, mod, resourcepack, script]),
        'dirs': {
            'config': (config, True),
            'mods': (mod, True),
            'resourcepack': (resourcepack, True),
            'script': (script, True)
        }
    }
}

def write_dir_entries(file_writer, path, enabled):
    if not path or not enabled:
        file_writer.write('{}')
        return
    files = getFiles(path)
    file_writer.write('{')
    first = True
    for name, file_path in files:
        if not first:
            file_writer.write(', ')
        # 写入键名
        file_writer.write(f'"{name}": ')
        # 写入hex数据
        hex_data = fromFileToBinary(file_path)
        if hex_data is None:
            file_writer.write('None')
        else:
            file_writer.write(f'"{hex_data}"')
        first = False
    file_writer.write('}')

# 开始写入文件
with open(output, "w", encoding='utf-8') as f:
    f.write('BIN = {\n')
    # 写入default部分
    f.write('    "default": {\n')
    for key, (path, enabled) in dirs['default'].items():
        f.write(f'        "{key}": ')
        write_dir_entries(f, path, enabled)
        f.write(',\n')
    f.write('    },\n')
    # 写入optional部分
    f.write(f'    "optional": ({dirs["optional"]["enabled"]}, {{\n')
    for key, (path, enabled) in dirs['optional']['dirs'].items():
        f.write(f'        "{key}": ')
        write_dir_entries(f, path, enabled)
        f.write(',\n')
    f.write('    })\n')
    f.write('}')

3. 额外优化建议

  • 如果最终是要还原文件,考虑用base64编码代替hex,编码后体积是原文件的1.33倍,比hex的2倍更节省内存和磁盘空间。
  • 对于不需要频繁读取的大文件,甚至可以考虑只存储文件哈希和路径,而不是整个二进制数据,按需读取原文件。

内容的提问来源于stack exchange,提问作者Berdy Alexei Cadaeib Fecei

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 16:14:58