Java中UTF-8解码失败:如何将HTML字符实体还原为日文?
Java还原HTML字符实体为日文文本
直接使用getBytes(UTF-8)无法解决这个问题,因为年这类是HTML对Unicode码点的转义字符串,不属于字节编码层面的问题,需要通过HTML实体解析逻辑来转换。以下是两种可行的解决方案:
方案1:使用Java内置工具类(无需第三方依赖)
借助Swing中的HTMLEditorKit和ParserDelegator来解析HTML实体:
import javax.swing.text.html.HTMLEditorKit; import javax.swing.text.html.parser.ParserDelegator; import java.io.StringReader; import java.io.IOException; public class HtmlEntityDecoder { public static String decodeHtmlEntities(String html) throws IOException { final StringBuilder sb = new StringBuilder(); HTMLEditorKit.ParserCallback callback = new HTMLEditorKit.ParserCallback() { @Override public void handleText(char[] data, int pos) { sb.append(data); } }; new ParserDelegator().parse(new StringReader(html), callback, false); return sb.toString(); } public static void main(String[] args) throws IOException { String htmlEntities = "年 ネン"; String decodedText = decodeHtmlEntities(htmlEntities); System.out.println(decodedText); // 输出:年 ネン } }
方案2:使用Apache Commons Text库(更简洁)
如果项目允许引入第三方依赖,使用StringEscapeUtils可以一行完成解析:
import org.apache.commons.text.StringEscapeUtils; public class HtmlEntityDecoder { public static void main(String[] args) { String htmlEntities = "年 ネン"; String decodedText = StringEscapeUtils.unescapeHtml4(htmlEntities); System.out.println(decodedText); // 输出:年 ネン } }
注意:使用Apache Commons Text时,需要在项目中引入对应依赖(比如Maven依赖:
commons-text:commons-text:2.10.0)
内容的提问来源于stack exchange,提问作者user273121
相关产品推荐
相关产品推荐

