You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 3:如何将ASCII字符转为Unicode转义序列以适配正则匹配

Solution: Convert ASCII Characters (via HTML Entities) to Unicode Escape Sequences for Regex Matching

Got it, let's fix this matching issue. The root problem is your search string uses HTML entities (like &) while the raw data stores the same character as a Unicode escape sequence (\u0026). To get them to match in regex, we need to convert the search string into the same escape format as the raw data.

Step-by-Step Approach

  1. Decode HTML Entities: First, turn HTML entities like & back into their original ASCII characters (in this case, &).
  2. Convert to Unicode Escape Sequences: Take each ASCII character and convert it to the \\uXXXX format (double backslash because we need to escape the backslash in Python strings for regex use).

Python Implementation

Here's a reusable function that handles both steps, plus regex-safe handling for special characters:

import html
import re

def convert_to_regex_safe_unicode_escape(search_str):
    # Step 1: Decode HTML entities to raw characters
    decoded = html.unescape(search_str)
    
    # Step 2: Convert each ASCII character to \\uXXXX escape sequence
    escaped_components = []
    for char in decoded:
        char_code = ord(char)
        # Handle ASCII characters (0-127) with Unicode escape
        if 0 <= char_code <= 127:
            escaped_components.append(f'\\\\u{char_code:04x}')
        else:
            # For non-ASCII, escape any regex special characters
            escaped_components.append(re.escape(char))
    
    return ''.join(escaped_components)

# Test with your example
teste = "teste's teste &amp; teste"
converted_pattern = convert_to_regex_safe_unicode_escape(teste)
print(converted_pattern)
# Output: \\u0074\\u0065\\u0073\\u0074\\u0065\\u0027\\u0073\\u0020\\u0074\\u0065\\u0073\\u0074\\u0065\\u0020\\u0026\\u0020\\u0074\\u0065\\u0073\\u0074\\u0065

# Now use this pattern to match your raw data
raw_data = '.... teste\'s teste \\u0026 teste",null,["here","here2"] ....'
match = re.search(f'{converted_pattern}(.*?)\["(.*?)","(.*?)"\]', raw_data)

if match:
    print("Matched items:", match.group(2), match.group(3))
# Output: Matched items: here here2

Key Details

  • HTML Entity Decoding: The html.unescape() function handles all standard HTML entities (not just &amp;), so this works for &lt;, &gt;, &quot;, etc.
  • Regex Safety: For non-ASCII characters, we use re.escape() to ensure characters like ., *, or ? don't get interpreted as regex special operators.
  • Unicode Escape Format: The \\uXXXX format matches exactly the escape sequences in your raw data (since raw data stores \\u0026 as a single backslash plus u0026 in the actual string).

内容的提问来源于stack exchange,提问作者Iury Fukuda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 04:27:35