You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:如何用Python过滤capa分析输出文本,提取目标特征列表

Capa输出特征提取过滤修正方案

问题场景

我使用capa分析恶意文件,将输出重定向到文本文件后,需要提取其中的唯一特征(例如decompress data using aPLib这类)存入新列表。已经在values列表中定义了不需要保留的内容,但现有代码无法正确过滤,请求指导。

示例输入(file_details内容)

['md5                     318271d61f4c88ccce4384ecea22807b', 
 'sha1                    2396ad920f108dad605e9d7664e38bae0ece685c', 
 'sha256                  3cc170a25027b48537aca361ed718e77ee86d028203e8a710e71f1136b5c674a', 
 'path                    /home/zakaria/code/malware/Backdoor.Win32.BlackHole.asva_e7e4.exe', 
 'timestamp               2023-09-19 11:47:28.203754', 
 'capa version            6.1.0', 
 'os                      windows', 
 'format                  pe', 
 'arch                    i386', 
 'extractor               VivisectFeatureExtractor', 
 'base address            0x400000', 
 'rules                   /tmp/_MEIQJyrc9/rules', 
 'function count          4', 
 'library function count  0', 
 'total feature count     21333', 
 '', 
 'decompress data using aPLib', 
 'namespace    data-manipulation/compression', 
 'description  detects decompression function of library aPLib', 
 'scope        function', 
 'matches      0x77470B', 
 '', '', '']

现有问题代码

import subprocess
import os
import csv

# 不需要保留的关键词/前缀
values = ['md5', 'sha1', 'sha256', 'path', 'timestamp', 'capa version', 'os', 'format', 'arch', 'extractor', 'base address', 'rules', 
          'function count', 'library function count', 'total feature count', 'namespace', 'scope', 'matches', '0x',
          '(internal) packer file limitation', 'Packed samples have often been obfuscated to hide their logic.',
          'capa cannot handle obfuscation well. This means the results may be misleading or incomplete.',
          '(internal) Visual Basic file limitation', 'description  This sample appears to be compiled from Visual Basic.',
          'representation called P-Code.', 'capa cannot handle Visual Basic executables well. This means that the results will be misleading or incomplete.',
          'You may have to analyze the file manually, for example using a tool like VB Decompiler.',
          'persist via Run registry key', 'packed with pebundle', 'capa cannot handle obfuscation well. This means the results may be misleading or incomplete.']

filename = 'Backdoor.Win32.BlackHole.asva_e7e4.exe' 
csv_file = '/home/zakaria/code/features.csv'
path_to_file = f'/home/zakaria/code/malware/{filename}'
result_list = [] # 存储需要保留的特征

with open("output.txt", "w") as output_file: 
    subprocess.run(['./capa', '-v', path_to_file], stdout=output_file) 

with open ('output.txt' , 'r') as raw_file:
    file_details = [line.strip() for line in raw_file]
for line in file_details:
    if line not in (value in line for value in values):
        result_list.append(line)

print(result_list)

问题原因

核心错误在过滤判断逻辑:if line not in (value in line for value in values),这个表达式会生成一堆布尔值(True/False),而line是字符串,永远不可能属于布尔值集合,导致所有行都被保留,完全起不到过滤作用。

修正方案

正确逻辑是:

  1. 跳过空行
  2. 检查当前行是否不包含values列表中的任何关键词/前缀,若不包含则保留

修正后的代码

import subprocess
import os
import csv

# 不需要保留的关键词/前缀
values = ['md5', 'sha1', 'sha256', 'path', 'timestamp', 'capa version', 'os', 'format', 'arch', 'extractor', 'base address', 'rules', 
          'function count', 'library function count', 'total feature count', 'namespace', 'scope', 'matches', '0x',
          '(internal) packer file limitation', 'Packed samples have often been obfuscated to hide their logic.',
          'capa cannot handle obfuscation well. This means the results may be misleading or incomplete.',
          '(internal) Visual Basic file limitation', 'description  This sample appears to be compiled from Visual Basic.',
          'representation called P-Code.', 'capa cannot handle Visual Basic executables well. This means that the results will be misleading or incomplete.',
          'You may have to analyze the file manually, for example using a tool like VB Decompiler.',
          'persist via Run registry key', 'packed with pebundle', 'capa cannot handle obfuscation well. This means the results may be misleading or incomplete.']

filename = 'Backdoor.Win32.BlackHole.asva_e7e4.exe' 
csv_file = '/home/zakaria/code/features.csv'
path_to_file = f'/home/zakaria/code/malware/{filename}'
result_list = [] # 存储需要保留的特征

# 运行capa并保存输出
with open("output.txt", "w") as output_file: 
    subprocess.run(['./capa', '-v', path_to_file], stdout=output_file, text=True) 

# 读取并过滤内容
with open('output.txt' , 'r') as raw_file:
    for line in raw_file:
        line_stripped = line.strip()
        # 跳过空行
        if not line_stripped:
            continue
        # 检查是否包含任何不需要的关键词,若都不包含则保留
        should_exclude = any(keyword in line_stripped for keyword in values)
        if not should_exclude:
            result_list.append(line_stripped)

print(result_list)

关键修正点说明

  1. 判断逻辑替换:用any(keyword in line_stripped for keyword in values)检查当前行是否包含任何不需要的关键词,返回True则排除,False则保留
  2. 空行处理:直接跳过空字符串,避免无效内容进入结果列表
  3. subprocess参数优化:添加text=True确保输出以文本模式写入,避免编码问题

测试结果

针对示例输入,运行修正后的代码会得到:

['decompress data using aPLib']

符合预期的特征提取需求。

内容的提问来源于stack exchange,提问作者JAKY

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 17:40:16