You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Java实现类似JavaScript decodeURIComponent(escape(x))的功能,恢复错误编码后的乱码字符串?

How to Recover the Original String from Garbled Text in Java (Equivalent to decodeURIComponent(escape(x)))

Let's break down what the JavaScript decodeURIComponent(escape(x)) does first, then map that logic to Java for a clean solution:

Understanding the JavaScript Logic

  • escape(x) takes the garbled string x (which comes from UTF-8 bytes incorrectly decoded as ISO-8859-1) and converts each character to its corresponding byte value in %XX format. Since ISO-8859-1 is a single-byte encoding, every character maps directly to one byte—exactly matching how escape handles non-ASCII characters here.
  • decodeURIComponent() then parses those %XX sequences back into UTF-8 bytes, and decodes them correctly into the original Unicode string.

Java Implementation Steps

To replicate this behavior in Java, follow these straightforward steps:

  1. Encode the garbled string to a byte array using ISO-8859-1—this reverses the incorrect decoding step, retrieving the original UTF-8 bytes.
  2. Convert each byte in the array to its URI-encoded %XX string representation (matching what escape() outputs).
  3. Decode this URI-encoded string using UTF-8 to get the original, correct string.

Full Code Example

import java.io.UnsupportedEncodingException;
import java.net.URLDecoder;
import java.net.URLEncoder;

public class GarbledStringRecovery {
    public static void main(String[] args) throws UnsupportedEncodingException {
        // The garbled string from your question
        String garbledText = "ÐÑибка валидаÑии аÑÑибÑÑов докÑменÑа";
        
        // Step 1: Get ISO-8859-1 bytes from the garbled string
        byte[] isoBytes = garbledText.getBytes("ISO-8859-1");
        
        // Step 2: Convert bytes to URI-encoded string (matches escape() behavior)
        String uriEncoded = URLEncoder.encode(new String(isoBytes, "ISO-8859-1"), "UTF-8");
        
        // Step 3: Decode URI string to recover original UTF-8 string
        String originalString = URLDecoder.decode(uriEncoded, "UTF-8");
        
        // Output the result
        System.out.println("Recovered original string: " + originalString);
        // Expected output: Ошибка валидации атрибутов документа
    }
}

Explanation

  • ISO-8859-1 Encoding: This is critical because it preserves the exact byte values of the garbled characters. Since the garbled text was created by misinterpreting UTF-8 bytes as ISO-8859-1, reversing this step gives us back the original UTF-8 byte sequence.
  • URI Encoding/Decoding: URLEncoder.encode() converts the ISO-8859-1 string to %XX format, just like escape(). URLDecoder.decode() then reads those %XX sequences as UTF-8 bytes and decodes them into the correct Unicode string.

Shortened One-Liner

For a more concise version (while keeping the same logic), you can combine all steps into a single line:

String original = URLDecoder.decode(URLEncoder.encode(garbledText, "ISO-8859-1"), "UTF-8");

This one-liner behaves exactly like decodeURIComponent(escape(x)) in JavaScript.

内容的提问来源于stack exchange,提问作者Eugene

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 04:44:05