如何使用Java实现类似JavaScript decodeURIComponent(escape(x))的功能,恢复错误编码后的乱码字符串?
How to Recover the Original String from Garbled Text in Java (Equivalent to
decodeURIComponent(escape(x))) Let's break down what the JavaScript decodeURIComponent(escape(x)) does first, then map that logic to Java for a clean solution:
Understanding the JavaScript Logic
escape(x)takes the garbled stringx(which comes from UTF-8 bytes incorrectly decoded as ISO-8859-1) and converts each character to its corresponding byte value in%XXformat. Since ISO-8859-1 is a single-byte encoding, every character maps directly to one byte—exactly matching howescapehandles non-ASCII characters here.decodeURIComponent()then parses those%XXsequences back into UTF-8 bytes, and decodes them correctly into the original Unicode string.
Java Implementation Steps
To replicate this behavior in Java, follow these straightforward steps:
- Encode the garbled string to a byte array using ISO-8859-1—this reverses the incorrect decoding step, retrieving the original UTF-8 bytes.
- Convert each byte in the array to its URI-encoded
%XXstring representation (matching whatescape()outputs). - Decode this URI-encoded string using UTF-8 to get the original, correct string.
Full Code Example
import java.io.UnsupportedEncodingException; import java.net.URLDecoder; import java.net.URLEncoder; public class GarbledStringRecovery { public static void main(String[] args) throws UnsupportedEncodingException { // The garbled string from your question String garbledText = "ÐÑибка валидаÑии аÑÑибÑÑов докÑменÑа"; // Step 1: Get ISO-8859-1 bytes from the garbled string byte[] isoBytes = garbledText.getBytes("ISO-8859-1"); // Step 2: Convert bytes to URI-encoded string (matches escape() behavior) String uriEncoded = URLEncoder.encode(new String(isoBytes, "ISO-8859-1"), "UTF-8"); // Step 3: Decode URI string to recover original UTF-8 string String originalString = URLDecoder.decode(uriEncoded, "UTF-8"); // Output the result System.out.println("Recovered original string: " + originalString); // Expected output: Ошибка валидации атрибутов документа } }
Explanation
- ISO-8859-1 Encoding: This is critical because it preserves the exact byte values of the garbled characters. Since the garbled text was created by misinterpreting UTF-8 bytes as ISO-8859-1, reversing this step gives us back the original UTF-8 byte sequence.
- URI Encoding/Decoding:
URLEncoder.encode()converts the ISO-8859-1 string to%XXformat, just likeescape().URLDecoder.decode()then reads those%XXsequences as UTF-8 bytes and decodes them into the correct Unicode string.
Shortened One-Liner
For a more concise version (while keeping the same logic), you can combine all steps into a single line:
String original = URLDecoder.decode(URLEncoder.encode(garbledText, "ISO-8859-1"), "UTF-8");
This one-liner behaves exactly like decodeURIComponent(escape(x)) in JavaScript.
内容的提问来源于stack exchange,提问作者Eugene
相关产品推荐
相关产品推荐

