如何通用提取不同哈希类型字符串中=后引号内的哈希值?
问题:提取哈希字符串中等号右侧引号内的哈希值
现有哈希字符串条目如下:
[file:hashes.'MD5' = '547334e75ed7d4eea2953675b07986b4'] [file:hashes.'SHA1' = '82d29b52e35e7938e7ee610c04ea9daaf5e08e90'] [file:hashes.'SHA256' = 'ff3b45ecfbbdb780b48b4c829d2b6078d8f7673d823bedbd6321699770fa3f84']
原本使用的Python脚本依赖固定索引,仅能提取MD5哈希值,无法适配哈希类型变化的情况:
if item['hash'][:12]=='[file:hashes': #从JSON字典中定位哈希字符串 if item['hash'][22:-2] not in hash_column: #仅能提取MD5的哈希值 insert_hash_table(item['hash'][22:-2]) #将哈希值插入表中
需要实现通用方法,提取所有这类字符串中等号右侧引号内的哈希值。
解决方案
方法1:字符串分割法
通过等号和单引号作为分割标记,分步提取哈希值,逻辑简单直观:
hash_str = item['hash'] if hash_str.startswith('[file:hashes'): # 按等号分割字符串,取右侧部分并去除首尾空格 right_part = hash_str.split('=')[1].strip() # 按单引号分割,取中间的哈希值内容 hash_value = right_part.split("'")[1] if hash_value not in hash_column: insert_hash_table(hash_value)
方法2:正则表达式法
用正则匹配等号右侧单引号包裹的十六进制哈希值,适配格式微小变化的场景:
import re hash_str = item['hash'] if hash_str.startswith('[file:hashes'): # 匹配等号后单引号内的十六进制字符(哈希值) match_result = re.search(r"= '([a-fA-F0-9]+)'", hash_str) if match_result: hash_value = match_result.group(1) if hash_value not in hash_column: insert_hash_table(hash_value)
内容的提问来源于stack exchange,提问作者Nima Sayyah
相关产品推荐
相关产品推荐

