You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于正则表达式解析可变长度带空格十六进制字符串的技术问询

Got it, let's work through this regex problem you're having with Python's re library. The issue with your original pattern is that it stops at the first space in the hex string—let's fix that and add some context to make it reliable.

The Fix: A Regex That Handles Space-Separated Hex Pairs

Your original regex doesn't account for the spaces between hex groups. Here's a pattern that will capture the full hex string, including all space-separated segments:

^(\w+)\s*:\s*((?:[A-Fa-f\d]{2}(?:\s+[A-Fa-f\d]{2})*))$

Let's break down what each part does:

  • ^ and $: Anchor the match to the start and end of a line (prevents matching partial lines or random text in the middle of a line)
  • (\w+): Captures the identifier (KEK, Key, Wrap) as group 1
  • \s*:\s*: Matches the colon, allowing any number of spaces (including zero) before/after it
  • ((?:[A-Fa-f\d]{2}(?:\s+[A-Fa-f\d]{2})*)): The core part that captures the full hex string:
    • [A-Fa-f\d]{2}: Matches a single 2-character hex pair
    • (?:\s+[A-Fa-f\d]{2})*: Matches zero or more instances of "one or more spaces + hex pair" (the ?: makes this a non-capturing group, so it doesn't clutter our results)

Python Code Implementation

Here's how to use this pattern in Python, including steps to clean the hex string (remove spaces) and convert it to a byte array if needed:

import re

# Your sample input text
tech_spec = """
【示例】使用192位KEK对7八位字节(octets)密钥数据进行加密:
KEK : 5840df6e29b02af1 ab493b705bf16ea1 ae8338f4dcc176a8
Key : 466f7250617369
Wrap : afbeb0f07dfbf541 9200f2ccb50bb24f
"""

# Compile the regex (optional but efficient for repeated use)
pattern = re.compile(r'^(\w+)\s*:\s*((?:[A-Fa-f\d]{2}(?:\s+[A-Fa-f\d]{2})*))$', re.MULTILINE)

# Find all matches in the text
matches = pattern.findall(tech_spec)

# Process each match
for key_type, hex_raw in matches:
    # Remove spaces to get a continuous hex string
    hex_clean = hex_raw.replace(' ', '')
    # Convert to a byte array (hex -> bytes)
    hex_bytes = bytes.fromhex(hex_clean)
    
    print(f"Extracted {key_type}:")
    print(f"  Raw with spaces: {hex_raw}")
    print(f"  Clean hex string: {hex_clean}")
    print(f"  Byte array: {hex_bytes}\n")

Why Online Tools vs Python re Behaved Differently

Most online regex tools enable the multiline mode by default, which makes ^ and $ match the start/end of each line instead of the entire string. In Python, you need to explicitly add the re.MULTILINE flag (like in the code above) or process each line individually to get the same behavior.

Bonus: Use Key Length Info for Validation

You mentioned the spec includes key length details (like "7 octets of key data"). You can use this to verify your extracted data is correct:

# Example: Validate Key length (7 octets = 14 hex characters)
expected_key_octets = 7
for key_type, hex_raw in matches:
    if key_type == "Key":
        hex_clean = hex_raw.replace(' ', '')
        if len(hex_clean) == expected_key_octets * 2:
            print(f"✓ Key length matches expected {expected_key_octets} octets")
        else:
            print(f"✗ Key length mismatch: expected {expected_key_octets*2} hex chars, got {len(hex_clean)}")

内容的提问来源于stack exchange,提问作者maxadamcamb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 13:03:15