Python爬虫中解析JS十六进制转义字符串为实际值的方法
\x Hex Escapes in Python Hey there! I totally get the frustration here—when you pull a string with \x hex escapes from a script tag during web scraping, Python doesn't automatically parse those escapes like it does when you directly assign the string in the terminal. Let's walk through two reliable ways to fix this in your code.
Method 1: Use unicode-escape Decoding
This is the most straightforward approach for basic cases. We'll first encode the raw string to bytes, then decode it using the unicode-escape codec, which will resolve the \x sequences into their actual characters.
Example Code:
# Raw string extracted from the div/script element (note the double backslashes) raw_escaped_str = 'aHR0cDovL3BsdC5hbmltZWhlYXZlbi5ldS\\x7crc3lkc2QvQl8tX1RoZV\\x7cCZWdpbm5pbmctLTEtLTE1MjAwNDQxMzcubXA0P3d3NXc0MQ==' # Decode the escaped sequence decoded_str = raw_escaped_str.encode('utf-8').decode('unicode-escape') # Print the result print(decoded_str) # Output: 'aHR0cDovL3BsdC5hbmltZWhlYXZlbi5ldS|rc3lkc2QvQl8tX1RoZV|CZWdpbm5pbmctLTEtLTE1MjAwNDQxMzcubXA0P3d3NXc0MQ=='
Method 2: Use ast.literal_eval (Safer for Complex Cases)
If your string might contain other special characters (like nested quotes) or you want a more robust solution, ast.literal_eval is the way to go. This function parses the string as a Python literal, just like when you assign it directly in the terminal—so it automatically handles all standard escape sequences safely.
Example Code:
import ast # Raw escaped string from scraping raw_escaped_str = 'aHR0cDovL3BsdC5hbmltZWhlYXZlbi5ldS\\x7crc3lkc2QvQl8tX1RoZV\\x7cCZWdpbm5pbmctLTEtLTE1MjAwNDQxMzcubXA0P3d3NXc0MQ==' # Parse the string as a Python literal decoded_str = ast.literal_eval(f'"{raw_escaped_str}"') print(decoded_str) # Output: 'aHR0cDovL3BsdC5hbmltZWhlYXZlbi5ldS|rc3lkc2QvQl8tX1RoZV|CZWdpbm5pbmctLTEtLTE1MjAwNDQxMzcubXA0P3d3NXc0MQ=='
Why This Works
When you paste the string directly into the Python terminal, the interpreter parses it as a string literal and resolves all escape sequences automatically. But when you scrape it from a webpage, you're getting a raw string where \x is just two regular characters (\ and x). The methods above force Python to process those escape sequences like it does for literal strings.
内容的提问来源于stack exchange,提问作者Xayden Rosario

