Python 3:如何将ASCII字符转为Unicode转义序列以适配正则匹配
Solution: Convert ASCII Characters (via HTML Entities) to Unicode Escape Sequences for Regex Matching
Got it, let's fix this matching issue. The root problem is your search string uses HTML entities (like &) while the raw data stores the same character as a Unicode escape sequence (\u0026). To get them to match in regex, we need to convert the search string into the same escape format as the raw data.
Step-by-Step Approach
- Decode HTML Entities: First, turn HTML entities like
&back into their original ASCII characters (in this case,&). - Convert to Unicode Escape Sequences: Take each ASCII character and convert it to the
\\uXXXXformat (double backslash because we need to escape the backslash in Python strings for regex use).
Python Implementation
Here's a reusable function that handles both steps, plus regex-safe handling for special characters:
import html import re def convert_to_regex_safe_unicode_escape(search_str): # Step 1: Decode HTML entities to raw characters decoded = html.unescape(search_str) # Step 2: Convert each ASCII character to \\uXXXX escape sequence escaped_components = [] for char in decoded: char_code = ord(char) # Handle ASCII characters (0-127) with Unicode escape if 0 <= char_code <= 127: escaped_components.append(f'\\\\u{char_code:04x}') else: # For non-ASCII, escape any regex special characters escaped_components.append(re.escape(char)) return ''.join(escaped_components) # Test with your example teste = "teste's teste & teste" converted_pattern = convert_to_regex_safe_unicode_escape(teste) print(converted_pattern) # Output: \\u0074\\u0065\\u0073\\u0074\\u0065\\u0027\\u0073\\u0020\\u0074\\u0065\\u0073\\u0074\\u0065\\u0020\\u0026\\u0020\\u0074\\u0065\\u0073\\u0074\\u0065 # Now use this pattern to match your raw data raw_data = '.... teste\'s teste \\u0026 teste",null,["here","here2"] ....' match = re.search(f'{converted_pattern}(.*?)\["(.*?)","(.*?)"\]', raw_data) if match: print("Matched items:", match.group(2), match.group(3)) # Output: Matched items: here here2
Key Details
- HTML Entity Decoding: The
html.unescape()function handles all standard HTML entities (not just&), so this works for<,>,", etc. - Regex Safety: For non-ASCII characters, we use
re.escape()to ensure characters like.,*, or?don't get interpreted as regex special operators. - Unicode Escape Format: The
\\uXXXXformat matches exactly the escape sequences in your raw data (since raw data stores\\u0026as a single backslash plusu0026in the actual string).
内容的提问来源于stack exchange,提问作者Iury Fukuda
相关产品推荐
相关产品推荐

