从二进制文件提取资源部分损坏,求解决方案及工具推荐
问题:二进制文件提取PNG部分损坏,修复后仍有异常
问题背景
我正在做一个从二进制文件提取图片、音频等资源的项目,用Python脚本提取PNG时,部分文件损坏。相关代码如下:
main.py
import os from src.bin2png import bin2png for root, dirs, files in os.walk('.\Bin'): for file in files: if not file.endswith('.bin'): continue print(f'Found file: {file}') bin2png(os.path.join(root, file), './png') print('DONE')
bin2png.py
import os import re import binascii def bin2png(file_path: str, save_path: str): counter = 0 png_pattern = re.compile(r'89504E47.*?49454E44AE426082', re.IGNORECASE) file_path = file_path.replace('\\', '/') file_name = file_path.split('/')[-1][:-4] save_path = os.path.join(save_path, file_name) if not os.path.exists(save_path): os.mkdir(save_path) with open(file_path, 'rb') as file: for match in png_pattern.findall(str(binascii.hexlify(file.read()))): with open('{}/{}.png'.format(save_path, counter), 'wb+') as image: image.write(binascii.unhexlify(match)) counter += 1 if counter != 0: print('Successfully extracted {} PNGs from file {}'.format(counter, file_name)) else: print('Failed to find PNGs in file {}'.format(file_name))
检查二进制文件时,发现一段疑似被混淆的PNG头:
A9 70 6E 67 2D 2A 3A 2A 20 20 20 2D 69 68 64 72 20 20 20 5F 20 20 CD 5F 26 75 39 D4 8C EF B0 3E 34 AC C3 23 A9 A3 4D 1D A0 FA 5D F8 32 EF 60 76 75 18 E2 88 D4 24
编写了hex_modifier.py尝试修复头:
def modify_hex_string(hex_string): hex_values = hex_string.split() modified_values = [] for hex_value in hex_values: int_value = int(hex_value, 16) modified_value = abs(0x20 - int_value) modified_value = max(0, min(modified_value, 0xFF)) modified_values.append(format(modified_value, '02X')) result_string = ' '.join(modified_values) return result_string hex_string = "A9 70 6E 67 2D 2A 3A 2A 20 20 20 2D 69 68 64 72 20 20 20 5F 20 20 CD 5F 26 75 39 D4 8C EF B0 3E 34 AC C3 23 A9 A3 4D 1D A0 FA 5D F8 32 EF 60 76 75 18 E2 88 D4 24" result = modify_hex_string(hex_string) print(f"Result: [{result}]")
输出修复后的头:
Result: [89 50 4E 47 0D 0A 1A 0A 00 00 00 0D 49 48 44 52 00 00 00 3F 00 00 AD 3F 06 55 19 B4 6C CF 90 1E 14 8C A3 03 89 83 2D 03 80 DA 3D D8 12 CF 40 56 55 08 C2 68 B4 04]
现在文件头正确,但仍有部分资源损坏,请问问题出在哪?如何修复?有哪些工具或参考内容推荐?
问题分析与修复方案
核心问题点
正则匹配的转码错误
原代码中str(binascii.hexlify(file.read()))会将十六进制字节流转为包含b'前缀和'后缀的字符串,导致正则匹配时可能捕获到无效字符,破坏提取的PNG数据完整性。混淆还原算法错误
观察混淆后的PNG头:A9 70 6E 67对应正常头89 50 4E 47,实际是每个字节加上了0x20,还原时应该用字节值 - 0x20,而不是abs(0x20 - 字节值)。后者会在字节值小于0x20时错误反转,导致数据损坏。仅修复文件头,未处理整个数据流
混淆是作用于整个PNG文件的,只修复头无法解决后续数据的损坏问题,需要对提取到的完整混淆数据做逆处理。
修复步骤
1. 修正二进制转码逻辑
在bin2png.py中,将十六进制转码改为:
hex_data = binascii.hexlify(file.read()).decode('ascii')
避免引入b'和'的干扰。
2. 修正混淆还原算法
替换错误的绝对值计算,改为正确的逆混淆逻辑:
def decode_obfuscated_png(data: bytes) -> bytes: decoded = bytearray() for b in data: # 每个字节减去0x20,负数则加0x100保证在0-255范围 decoded_byte = (b - 0x20) % 0x100 decoded.append(decoded_byte) return bytes(decoded)
3. 优化正则匹配,处理混淆与原始PNG
更新正则表达式,同时匹配原始PNG和混淆后的PNG,并对混淆数据做还原:
png_pattern = re.compile( r'(89504E47.*?0000000049454E44AE426082)|(A9706E67.*?2020202069686472CE6280A2)', re.IGNORECASE | re.DOTALL )
- 原始PNG的IEND块为
0000000049454E44AE426082(长度0 + IEND签名 + CRC) - 混淆后的IEND块为
2020202069686472CE6280A2(每个字节加0x20后的结果)
4. 完整改进后的bin2png.py
import os import re import binascii def decode_obfuscated_png(data: bytes) -> bytes: decoded = bytearray() for b in data: decoded_byte = (b - 0x20) % 0x100 decoded.append(decoded_byte) return bytes(decoded) def bin2png(file_path: str, save_path: str): counter = 0 png_pattern = re.compile( r'(89504E47.*?0000000049454E44AE426082)|(A9706E67.*?2020202069686472CE6280A2)', re.IGNORECASE | re.DOTALL ) file_path = file_path.replace('\\', '/') file_name = file_path.split('/')[-1][:-4] save_path = os.path.join(save_path, file_name) if not os.path.exists(save_path): os.mkdir(save_path) with open(file_path, 'rb') as file: hex_data = binascii.hexlify(file.read()).decode('ascii') for match in png_pattern.findall(hex_data): png_hex = match[0] if match[0] else match[1] png_data = binascii.unhexlify(png_hex) # 对混淆的PNG数据做解码 if png_hex.lower().startswith('a9706e67'): png_data = decode_obfuscated_png(png_data) # 保存文件 with open(os.path.join(save_path, f'{counter}.png'), 'wb') as image: image.write(png_data) counter += 1 if counter != 0: print(f'Successfully extracted {counter} PNGs from file {file_name}') else: print(f'Failed to find PNGs in file {file_name}')
其他可能的问题排查
- PNG块CRC校验失败:如果提取的PNG仍损坏,可以用
png库检查每个块的CRC,验证数据是否被篡改或截断。 - 多混淆模式:部分PNG可能用了不同的混淆规则,需要用十六进制编辑器对比正常PNG和损坏PNG的字节差异,确定其他混淆方式。
推荐工具与参考
- 十六进制编辑器:如HxD,用于手动对比正常PNG与损坏PNG的字节结构,分析混淆规则。
- 二进制提取工具:foremost、binwalk,自动识别多种文件签名,支持批量提取,可作为脚本的补充。
- PNG格式文档:深入了解PNG的块结构(IHDR、IDAT、IEND等),帮助精准匹配和验证数据。
- Python库:
png库用于解析和验证PNG文件,检查块完整性与CRC。
内容的提问来源于stack exchange,提问作者Leo A. A. Ron
相关产品推荐
相关产品推荐

