You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

HTML ASCII代码‡无法转换为目标字符‡的问题求助

HTML转义字符�转换后为空的原因及解决方法

问题场景

输入字符串:"Hello, world! ‡ end"
期望输出:"Hello, world! ‡ end"
实际输出:"Hello, world! end"

已尝试两种方法均得到相同结果:

尝试1:使用Apache Commons Text

import org.apache.commons.text.StringEscapeUtils;
String str = "Hello, world! ‡ end";
String res = StringEscapeUtils.unescapeHtml4(str); // 用unescapeHtml3结果一致
System.out.println(res); // 输出:Hello, world!  end

尝试2:正则匹配转换

String getUpdatedStr() {
    String str = "Hello, world! ‡ end";
    String regex = "&#(\\d+);";
    Pattern pattern = Pattern.compile(regex);
    Matcher matcher = pattern.matcher(str);
    StringBuffer sb = new StringBuffer();
    while (matcher.find()) {
        int asciiValue = Integer.parseInt(matcher.group(1));
        char asciiChar = (char) asciiValue;
        matcher.appendReplacement(sb, Character.toString(asciiChar));
    }
    matcher.appendTail(sb);
    return sb.toString(); // 输出:Hello, world!  end
}

原因分析

  1. 标准ASCII与扩展ASCII的混淆
    标准ASCII仅覆盖0-127的字符,‡属于**扩展ASCII(如Windows-1252编码)**范畴。但Java的char类型基于Unicode编码,直接将十进制135转换为char得到的是Unicode U+0087——这是一个C1控制字符,无可见显示符号,因此输出时呈现为空。

  2. 字符映射错误
    你参考的ASCII表中135对应的‡(双剑号),实际是Windows-1252编码下的扩展字符,其对应的Unicode编码为U+2021(十进制值为8225),而非135。直接用135转换自然无法得到目标可见字符。


解决方法

方法1:修正目标字符的Unicode映射

直接将‡替换为对应‡的Unicode值:

String getUpdatedStr() {
    String str = "Hello, world! ‡ end";
    String regex = "&#(\\d+);";
    Pattern pattern = Pattern.compile(regex);
    Matcher matcher = pattern.matcher(str);
    StringBuffer sb = new StringBuffer();
    while (matcher.find()) {
        int num = Integer.parseInt(matcher.group(1));
        char targetChar = num == 135 ? '\u2021' : (char) num;
        matcher.appendReplacement(sb, Character.toString(targetChar));
    }
    matcher.appendTail(sb);
    return sb.toString();
}

方法2:按指定编码处理扩展ASCII

批量处理扩展ASCII字符时,可将数值转为字节后按Windows-1252编码解码:

String getUpdatedStr() {
    String str = "Hello, world! ‡ end";
    String regex = "&#(\\d+);";
    Pattern pattern = Pattern.compile(regex);
    Matcher matcher = pattern.matcher(str);
    StringBuffer sb = new StringBuffer();
    while (matcher.find()) {
        int num = Integer.parseInt(matcher.group(1));
        String charStr;
        if (num > 127) {
            // 扩展ASCII按Windows-1252编码转换
            try {
                charStr = new String(new byte[]{(byte) num}, "Windows-1252");
            } catch (UnsupportedEncodingException e) {
                charStr = "";
                e.printStackTrace();
            }
        } else {
            charStr = Character.toString((char) num);
        }
        matcher.appendReplacement(sb, charStr);
    }
    matcher.appendTail(sb);
    return sb.toString();
}

方法3:修正HTML实体后再解码

若使用StringEscapeUtils,可先将‡替换为正确的Unicode实体”,再执行解码:

import org.apache.commons.text.StringEscapeUtils;
String str = "Hello, world! ‡ end";
// 替换为正确的HTML实体
String correctedStr = str.replace("‡", "”");
String res = StringEscapeUtils.unescapeHtml4(correctedStr);
System.out.println(res); // 输出:Hello, world! ‡ end

内容的提问来源于stack exchange,提问作者kar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 06:53:17