HTML ASCII代码‡无法转换为目标字符‡的问题求助
HTML转义字符�转换后为空的原因及解决方法
问题场景
输入字符串:"Hello, world! ‡ end"
期望输出:"Hello, world! ‡ end"
实际输出:"Hello, world! end"
已尝试两种方法均得到相同结果:
尝试1:使用Apache Commons Text
import org.apache.commons.text.StringEscapeUtils; String str = "Hello, world! ‡ end"; String res = StringEscapeUtils.unescapeHtml4(str); // 用unescapeHtml3结果一致 System.out.println(res); // 输出:Hello, world! end
尝试2:正则匹配转换
String getUpdatedStr() { String str = "Hello, world! ‡ end"; String regex = "&#(\\d+);"; Pattern pattern = Pattern.compile(regex); Matcher matcher = pattern.matcher(str); StringBuffer sb = new StringBuffer(); while (matcher.find()) { int asciiValue = Integer.parseInt(matcher.group(1)); char asciiChar = (char) asciiValue; matcher.appendReplacement(sb, Character.toString(asciiChar)); } matcher.appendTail(sb); return sb.toString(); // 输出:Hello, world! end }
原因分析
标准ASCII与扩展ASCII的混淆
标准ASCII仅覆盖0-127的字符,‡属于**扩展ASCII(如Windows-1252编码)**范畴。但Java的char类型基于Unicode编码,直接将十进制135转换为char得到的是Unicode U+0087——这是一个C1控制字符,无可见显示符号,因此输出时呈现为空。字符映射错误
你参考的ASCII表中135对应的‡(双剑号),实际是Windows-1252编码下的扩展字符,其对应的Unicode编码为U+2021(十进制值为8225),而非135。直接用135转换自然无法得到目标可见字符。
解决方法
方法1:修正目标字符的Unicode映射
直接将‡替换为对应‡的Unicode值:
String getUpdatedStr() { String str = "Hello, world! ‡ end"; String regex = "&#(\\d+);"; Pattern pattern = Pattern.compile(regex); Matcher matcher = pattern.matcher(str); StringBuffer sb = new StringBuffer(); while (matcher.find()) { int num = Integer.parseInt(matcher.group(1)); char targetChar = num == 135 ? '\u2021' : (char) num; matcher.appendReplacement(sb, Character.toString(targetChar)); } matcher.appendTail(sb); return sb.toString(); }
方法2:按指定编码处理扩展ASCII
批量处理扩展ASCII字符时,可将数值转为字节后按Windows-1252编码解码:
String getUpdatedStr() { String str = "Hello, world! ‡ end"; String regex = "&#(\\d+);"; Pattern pattern = Pattern.compile(regex); Matcher matcher = pattern.matcher(str); StringBuffer sb = new StringBuffer(); while (matcher.find()) { int num = Integer.parseInt(matcher.group(1)); String charStr; if (num > 127) { // 扩展ASCII按Windows-1252编码转换 try { charStr = new String(new byte[]{(byte) num}, "Windows-1252"); } catch (UnsupportedEncodingException e) { charStr = ""; e.printStackTrace(); } } else { charStr = Character.toString((char) num); } matcher.appendReplacement(sb, charStr); } matcher.appendTail(sb); return sb.toString(); }
方法3:修正HTML实体后再解码
若使用StringEscapeUtils,可先将‡替换为正确的Unicode实体”,再执行解码:
import org.apache.commons.text.StringEscapeUtils; String str = "Hello, world! ‡ end"; // 替换为正确的HTML实体 String correctedStr = str.replace("‡", "”"); String res = StringEscapeUtils.unescapeHtml4(correctedStr); System.out.println(res); // 输出:Hello, world! ‡ end
内容的提问来源于stack exchange,提问作者kar
相关产品推荐
相关产品推荐

