You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java中如何移除字符串的Unicode杂项符号十进制表示

Fixing Unicode Entity Conversion/Removal in Java

First, let's figure out why your initial regex didn't work: your pattern &#[9728 - 9983]; uses a character class [] incorrectly — it's matching individual characters (like 9,7,2,8, space, -, etc.) instead of the numeric range you intended. Let's cover both solutions you asked for: converting those entities back to their original emoji symbols, and removing them entirely.

Solution 1: Convert Unicode HTML Entities Back to Original Emojis

If you want to preserve the original symbols (⛔, ⚡, ⚠) instead of deleting them, you have two straightforward options:

Option A: Use Apache Commons Text (Easiest)

The StringEscapeUtils class from Apache Commons Text handles HTML entity unescaping out of the box. It can resolve both &#xxxx; (double-escaped entities) and plain &#xxxx; entities:

import org.apache.commons.text.StringEscapeUtils;

public class EmojiConverter {
    public static void main(String[] args) {
        String input = "Last Search Results Bulletin Board Validations ⛔ 0 ⚡ 1 ⚠ 6 ? 0";
        
        // Unescape twice: first to turn & into &, then to resolve &#xxxx; to emoji
        String output = StringEscapeUtils.unescapeHtml4(StringEscapeUtils.unescapeHtml4(input));
        
        System.out.println(output);
        // Output: Last Search Results Bulletin Board Validations ⛔ 0 ⚡ 1 ⚠ 6 ? 0
    }
}

Option B: Custom Regex (No Third-Party Libraries)

If you don't want to add a dependency, use Java's regex with a matcher to convert the numeric entities to characters. This works for Java 9+ (using lambda in replaceAll):

public class EmojiConverter {
    public static void main(String[] args) {
        String input = "Last Search Results Bulletin Board Validations ⛔ 0 ⚡ 1 ⚠ 6 ? 0";
        
        // Handle double-escaped entities (&#xxxx;) first
        String result = input.replaceAll("&#(\\d+);", matcher -> {
            int codePoint = Integer.parseInt(matcher.group(1));
            return String.valueOf(Character.toChars(codePoint));
        });
        
        // Handle any remaining plain &#xxxx; entities
        result = result.replaceAll("&#(\\d+);", matcher -> {
            int codePoint = Integer.parseInt(matcher.group(1));
            return String.valueOf(Character.toChars(codePoint));
        });
        
        System.out.println(result);
    }
}

For Java 8 or older, use Matcher.appendReplacement:

import java.util.regex.Matcher;
import java.util.regex.Pattern;

public class EmojiConverter {
    public static void main(String[] args) {
        String input = "Last Search Results Bulletin Board Validations ⛔ 0 ⚡ 1 ⚠ 6 ? 0";
        Pattern pattern = Pattern.compile("&#(\\d+);|&#(\\d+);");
        Matcher matcher = pattern.matcher(input);
        StringBuffer sb = new StringBuffer();
        
        while (matcher.find()) {
            // Grab the numeric group (either group 1 or 2)
            String numStr = matcher.group(1) != null ? matcher.group(1) : matcher.group(2);
            int codePoint = Integer.parseInt(numStr);
            matcher.appendReplacement(sb, String.valueOf(Character.toChars(codePoint)));
        }
        matcher.appendTail(sb);
        
        System.out.println(sb.toString());
    }
}

Solution 2: Remove the Unicode Entities Entirely

If you just want to delete these entities, use a regex that targets the exact entity formats.

Remove All Such Entities

This will delete any &#xxxx; or &#xxxx; pattern:

public class EntityRemover {
    public static void main(String[] args) {
        String input = "Last Search Results Bulletin Board Validations ⛔ 0 ⚡ 1 ⚠ 6 ? 0";
        
        String cleaned = input.replaceAll("&#\\d+;", "")
                              .replaceAll("&#\\d+;", "");
        
        System.out.println(cleaned);
        // Output: Last Search Results Bulletin Board Validations  0  1  6 ? 0
    }
}

Remove Only Entities in Your Target Range (9728-9983)

If you want to delete only entities in the emoji range you specified, use a regex that matches that numeric range specifically:

public class EntityRemover {
    public static void main(String[] args) {
        String input = "Last Search Results Bulletin Board Validations ⛔ 0 ⚡ 1 ⚠ 6 ? 0";
        
        // Regex to match 9728-9983: covers all values in your range
        String rangeRegex = "(972[8-9]|97[3-9]\\d|9[8-9]\\d{2}|998[0-3])";
        
        String cleaned = input.replaceAll("&#" + rangeRegex + ";", "")
                              .replaceAll("&#" + rangeRegex + ";", "");
        
        System.out.println(cleaned);
    }
}

Bonus: Prevent This Issue at the Source

To avoid having emojis converted to entities in the first place, make sure:

  1. Your database and table columns use the utf8mb4 character set (supports full Unicode, including 4-byte emojis).
  2. Your web form doesn't automatically escape HTML entities for the text field (or unescape them before saving to the database).

内容的提问来源于stack exchange,提问作者nagaraju

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:43:47