You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将十六进制原始字节转换为RGB图像(恶意软件检测数据集)

恶意软件十六进制字节转RGB图像实现

我正在通过机器学习探索恶意软件检测技术,偶然发现Kaggle上的微软2015大型数据集。该数据集包含十六进制格式的字节数据,示例如下:

00401000 00 00 80 40 40 28 00 1C 02 42 00 C4 00 20 04 20
00401010 00 00 20 09 2A 02 00 00 00 00 8E 10 41 0A 21 01
00401020 40 00 02 01 00 90 21 00 32 40 00 1C 01 40 C8 18
00582FF0 ?? ?? ?? ?? ?? ?? ?? ?? ?? ?? ?? ?? ?? ?? ?? ??

我已有一段将字节转换为灰度图像的脚本,现在需要基于它实现RGB图像的转换。

实现思路

灰度图像是单通道(每个像素用1字节表示亮度),而RGB图像是三通道(每个像素用3字节分别表示红、绿、蓝分量)。转换核心逻辑:

  • 收集所有有效字节(将??替换为0),整理为一维数组
  • 按每3个字节一组分配给R、G、B通道,剩余不足3个的字节补0填充
  • 计算合适的图像尺寸(沿用原脚本中2的幂次尺寸逻辑,保证图像规整)
  • 将三通道数据组合为RGB格式数组,保存为图像

修改后的Python代码

import os
from math import log
import numpy as np
from PIL import Image

def save_rgb_img(byte_array, name):
    # 将一维字节数组转换为RGB三通道格式
    total_bytes = len(byte_array)
    # 补全到3的倍数,不足的补0
    pad_length = (3 - total_bytes % 3) % 3
    byte_array = np.pad(byte_array, (0, pad_length), mode='constant', constant_values=0)
    
    # 拆分R、G、B通道
    r_channel = byte_array[0::3]
    g_channel = byte_array[1::3]
    b_channel = byte_array[2::3]
    
    # 计算图像尺寸:总像素数为 len(byte_array)//3,找最接近的2的幂次宽高
    total_pixels = len(r_channel)
    side = int(total_pixels ** 0.5)
    # 取大于等于side的最小2的幂次作为宽或高
    side = 2 ** (int(log(side)/log(2)) + 1)
    # 计算另一个维度,确保总像素数不超过实际像素数
    height = int(total_pixels / side)
    # 如果还有剩余像素,高度加1,补0填充
    if total_pixels % side != 0:
        height += 1
        pad_pixels = side * height - total_pixels
        r_channel = np.pad(r_channel, (0, pad_pixels), mode='constant')
        g_channel = np.pad(g_channel, (0, pad_pixels), mode='constant')
        b_channel = np.pad(b_channel, (0, pad_pixels), mode='constant')
    
    # 转换为二维通道数组
    r_channel = r_channel.reshape((height, side))
    g_channel = g_channel.reshape((height, side))
    b_channel = b_channel.reshape((height, side))
    
    # 合并为RGB三维数组
    rgb_array = np.stack([r_channel, g_channel, b_channel], axis=-1)
    
    # 保存图像
    im = Image.fromarray(np.uint8(rgb_array))
    im.save(f"{name}_rgb.jpg", "JPEG")
    print(f"RGB图像已保存:{name}_rgb.jpg,尺寸:{height}x{side}")

# 替换为你的字节文件目录路径
bytes_folder = "/test"
files = os.listdir(bytes_folder)

for filename in files:
    if not filename.endswith('.bytes'):
        continue
    
    file_path = os.path.join(bytes_folder, filename)
    byte_list = []
    
    with open(file_path, 'r') as f:
        for line in f:
            parts = line.split()
            # 每行应该是1个地址+16个字节,共17个元素
            if len(parts) != 17:
                continue
            # 提取16个字节,替换??为0并转为整数
            for byte_str in parts[1:]:
                byte_val = int(byte_str, 16) if byte_str != '??' else 0
                byte_list.append(byte_val)
    
    if byte_list:
        save_rgb_img(np.array(byte_list), filename[:-6])

代码说明

  • 数据读取:遍历.bytes文件,提取每行的16个字节,将??替换为0后收集到一维数组
  • 通道拆分:将一维数组按索引步长3拆分到R、G、B三个通道,不足的补0
  • 尺寸计算:沿用原脚本的2的幂次尺寸逻辑,保证图像为规整的矩形,避免拉伸变形
  • 图像保存:将三通道数组合并为RGB格式,用PIL保存为JPEG图像

内容的提问来源于stack exchange,提问作者benjamin olise

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 17:21:07