You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于表头对CSV恶意软件分析数据实现独热编码

恶意软件特征独热编码优化方案

数据背景

我手头有大量恶意软件分析数据,已整理出包含所有唯一特征的表头:

label,md5,create new key via CryptAcquireContext,delete user account group,send file via HTTP

CSV数据示例如下:

malware,0e498d2283a99aad9a997200caf2f24e,packed with UPX,delete user account group
malware,1f0ffe101843e4135af712d899b0c0d3,inspect load icon resource,reference analysis tools strings,send file via HTTP

需求

实现独热编码规则:哈希值(md5列)后的内容若与表头中的特征匹配则设为1,不匹配则设为0。

现有代码

import csv

def one_hot_encode(header, data):
    encoded_data = []
    for row in data:
        encoded_row = []

        # 遍历行内每个元素
        for element in row:
            # 检查元素是否匹配当前表头项
            if element == header:
                encoded_row.append(1)  # 匹配设为1
            else:
                encoded_row.append(0)  # 不匹配设为0

        encoded_data.append(encoded_row)

    return encoded_data

def main():
    headers = []
    output_file = "feature-list.txt"

    # 从文本文件读取表头
    with open(output_file, 'r') as txt_file:
        for line in txt_file:
            line = line.strip()  # 移除换行符
            headers.append(line)

    headers.insert(1, "md5")
    headers.insert(0, "label")
    input_file = 'cleaned_features-test-2.csv'
    output_file = 'cleaned_features-test2_updated.csv'

    # 读取CSV中的现有数据
    existing_data = []
    with open(input_file, 'r', newline='') as csv_file:
        reader = csv.reader(csv_file)
        for row in reader:
            existing_data.append(row)

    # 对每个表头项执行独热编码
    encoded_data = []
    for header in headers:
        encoded_data.extend(one_hot_encode(header, existing_data))

    # 将自定义表头和编码后数据写入新CSV
    with open(output_file, 'w', newline='') as csv_file:
        writer = csv.writer(csv_file)

        # 写入表头
        writer.writerow(headers)

        # 写入编码数据
        writer.writerows(encoded_data)

if __name__ == "__main__":
    main()

优化实现方案

现有代码存在逻辑偏差:当前实现会为每个表头项生成一整行编码,最终输出行数为原始数据行数×表头数,不符合独热编码的预期格式(应为原始每行对应一行编码,每列对应一个特征的0/1值)。以下是修正后的高效实现:

优化思路

  1. 拆分表头,分离出label、md5和特征列
  2. 用集合存储每行的特征,将存在性检查的时间复杂度从O(n)降至O(1)
  3. 边读边写数据,避免加载全部数据到内存,适配大文件场景
  4. 简化编码逻辑,直接为每行生成对应的编码行

优化代码

import csv

def main():
    # 读取表头文件(假设文件内为逗号分隔的完整表头)
    header_path = "feature-list.txt"
    with open(header_path, 'r') as f:
        headers = f.readline().strip().split(',')
    
    # 拆分表头:标签、md5、特征列
    label_col, md5_col, *feature_cols = headers
    
    input_path = 'cleaned_features-test-2.csv'
    output_path = 'cleaned_features-test2_updated.csv'

    with open(input_path, 'r', newline='') as infile, open(output_path, 'w', newline='') as outfile:
        reader = csv.reader(infile)
        writer = csv.writer(outfile)
        
        # 写入表头
        writer.writerow(headers)
        
        for row in reader:
            # 提取当前行的标签、md5和特征集合
            current_label, current_md5, *current_features = row
            feature_set = set(current_features)
            
            # 生成编码行:标签 + md5 + 每个特征的0/1值
            encoded_row = [current_label, current_md5]
            for feature in feature_cols:
                encoded_row.append(1 if feature in feature_set else 0)
            
            writer.writerow(encoded_row)

if __name__ == "__main__":
    main()

优化说明

  • 逻辑修正:每行原始数据对应一行编码,符合独热编码的标准输出格式
  • 效率提升:集合查询的时间复杂度更低,处理大规模数据时性能优势明显
  • 内存优化:无需将全部数据加载到内存,适合处理超大CSV文件
  • 代码简化:去掉冗余函数,逻辑更直观,便于后续维护

输出示例

针对提供的CSV示例,优化后的代码会输出:

label,md5,create new key via CryptAcquireContext,delete user account group,send file via HTTP
malware,0e498d2283a99aad9a997200caf2f24e,0,1,0
malware,1f0ffe101843e4135af712d899b0c0d3,0,0,1

内容的提问来源于stack exchange,提问作者thanks_pop

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 08:18:23