Python处理Minecraft模组大文件转二进制时触发Memory Error
解决Minecraft整合包大文件二进制转储时的MemoryError问题
我正在开发一个程序,用来为Minecraft整合包生成包含各类文件二进制数据的字典。小文件处理正常,但遇到90MB的“Better End Reforked”这类大模组文件时,程序直接报MemoryError崩溃。
我的代码:
from os import listdir from os.path import abspath, splitext, basename, getsize, isfile, join, expandvars def getBool(statement): return (True if statement else False) def getFiles(path): list = [] for file in listdir(path): filePath = join(path, file) if isfile(filePath): list.append((file, abspath(filePath))) return list def getIdentity(path): name = basename(path) return { 'name': name, 'basename': splitext(name)[0], 'extension': splitext(name)[1], 'realpath': abspath(expandvars(path)), 'size': getsize(path) } def fromFileToBinary(path): try: with open(path, 'rb') as file: binary = file.read() return binary.hex() except IOError: return None output = getIdentity('./output/bin.py')['realpath'] config = 'C:/Users/Berdy Alexei/Downloads/modpack/optional/config' mod = 'C:/Users/Berdy Alexei/Downloads/modpack/optional/mods' resourcepack = 'C:/Users/Berdy Alexei/Downloads/modpack/optional/resourcepacks' script = 'C:/Users/Berdy Alexei/Downloads/modpack/optional/scripts' bool = getBool(config or mod or resourcepack or script) def _bin(path = None, bool = True): dictionary = {} if path and bool: files = getFiles(path) for file in files: name, path = file dictionary[name] = fromFileToBinary(path) return dictionary content = { 'default': { 'config': _bin('C:/Users/Berdy Alexei/Downloads/modpack/default/config', False), 'mods': _bin('C:/Users/Berdy Alexei/Downloads/modpack/default/mods', False), 'resourcepack': _bin('C:/Users/Berdy Alexei/Downloads/modpack/default/resourcepacks', False), 'script': _bin('C:/Users/Berdy Alexei/Downloads/modpack/default/scripts', False) }, 'optional': (bool, { 'config': _bin(config), 'mods': _bin(mod), 'resourcepack': _bin(resourcepack), 'script': _bin(script) }) } with open(output, "w") as file: file.write('BIN = {}'.format(content))
报错信息:
Traceback (most recent call last): File "c:\Folders\Archivos\Proyectos\InstallerCrafter\InstallerCrafter.py", line 75, in <module> file.write('BIN = {}'.format(content)) MemoryError
处理的文件:
问题根源
- 一次性读取大文件:
fromFileToBinary里用file.read()直接把90MB的文件全读进内存,转成hex字符串后体积会翻倍(每个字节转成两个字符),光是这一个文件就占180MB内存。 - 全量存储再写入:所有文件的hex字符串都存在
content字典里,最后还要把整个字典转成一个巨大的字符串写入,内存直接被撑爆。
修复方案
1. 分块读取大文件转hex
修改fromFileToBinary,分块读取文件,避免一次性加载整个文件:
def fromFileToBinary(path): try: hex_str = [] with open(path, 'rb') as file: # 每次读4KB块,平衡效率和内存占用 chunk = file.read(4096) while chunk: hex_str.append(chunk.hex()) chunk = file.read(4096) return ''.join(hex_str) except IOError: return None
2. 边处理边写入文件,避免全量存内存
不要先把所有数据塞进字典再写入,而是直接生成Python代码并逐段写入文件,这样内存只会保留当前处理的文件数据:
output = getIdentity('./output/bin.py')['realpath'] # 定义要处理的目录 dirs = { 'default': { 'config': ('C:/Users/Berdy Alexei/Downloads/modpack/default/config', False), 'mods': ('C:/Users/Berdy Alexei/Downloads/modpack/default/mods', False), 'resourcepack': ('C:/Users/Berdy Alexei/Downloads/modpack/default/resourcepacks', False), 'script': ('C:/Users/Berdy Alexei/Downloads/modpack/default/scripts', False) }, 'optional': { 'enabled': any([config, mod, resourcepack, script]), 'dirs': { 'config': (config, True), 'mods': (mod, True), 'resourcepack': (resourcepack, True), 'script': (script, True) } } } def write_dir_entries(file_writer, path, enabled): if not path or not enabled: file_writer.write('{}') return files = getFiles(path) file_writer.write('{') first = True for name, file_path in files: if not first: file_writer.write(', ') # 写入键名 file_writer.write(f'"{name}": ') # 写入hex数据 hex_data = fromFileToBinary(file_path) if hex_data is None: file_writer.write('None') else: file_writer.write(f'"{hex_data}"') first = False file_writer.write('}') # 开始写入文件 with open(output, "w", encoding='utf-8') as f: f.write('BIN = {\n') # 写入default部分 f.write(' "default": {\n') for key, (path, enabled) in dirs['default'].items(): f.write(f' "{key}": ') write_dir_entries(f, path, enabled) f.write(',\n') f.write(' },\n') # 写入optional部分 f.write(f' "optional": ({dirs["optional"]["enabled"]}, {{\n') for key, (path, enabled) in dirs['optional']['dirs'].items(): f.write(f' "{key}": ') write_dir_entries(f, path, enabled) f.write(',\n') f.write(' })\n') f.write('}')
3. 额外优化建议
- 如果最终是要还原文件,考虑用base64编码代替hex,编码后体积是原文件的1.33倍,比hex的2倍更节省内存和磁盘空间。
- 对于不需要频繁读取的大文件,甚至可以考虑只存储文件哈希和路径,而不是整个二进制数据,按需读取原文件。
内容的提问来源于stack exchange,提问作者Berdy Alexei Cadaeib Fecei
相关产品推荐
相关产品推荐

