Java中如何移除字符串的Unicode杂项符号十进制表示
First, let's figure out why your initial regex didn't work: your pattern &#[9728 - 9983]; uses a character class [] incorrectly — it's matching individual characters (like 9,7,2,8, space, -, etc.) instead of the numeric range you intended. Let's cover both solutions you asked for: converting those entities back to their original emoji symbols, and removing them entirely.
Solution 1: Convert Unicode HTML Entities Back to Original Emojis
If you want to preserve the original symbols (⛔, ⚡, ⚠) instead of deleting them, you have two straightforward options:
Option A: Use Apache Commons Text (Easiest)
The StringEscapeUtils class from Apache Commons Text handles HTML entity unescaping out of the box. It can resolve both &#xxxx; (double-escaped entities) and plain &#xxxx; entities:
import org.apache.commons.text.StringEscapeUtils; public class EmojiConverter { public static void main(String[] args) { String input = "Last Search Results Bulletin Board Validations ⛔ 0 ⚡ 1 ⚠ 6 ? 0"; // Unescape twice: first to turn & into &, then to resolve &#xxxx; to emoji String output = StringEscapeUtils.unescapeHtml4(StringEscapeUtils.unescapeHtml4(input)); System.out.println(output); // Output: Last Search Results Bulletin Board Validations ⛔ 0 ⚡ 1 ⚠ 6 ? 0 } }
Option B: Custom Regex (No Third-Party Libraries)
If you don't want to add a dependency, use Java's regex with a matcher to convert the numeric entities to characters. This works for Java 9+ (using lambda in replaceAll):
public class EmojiConverter { public static void main(String[] args) { String input = "Last Search Results Bulletin Board Validations ⛔ 0 ⚡ 1 ⚠ 6 ? 0"; // Handle double-escaped entities (&#xxxx;) first String result = input.replaceAll("&#(\\d+);", matcher -> { int codePoint = Integer.parseInt(matcher.group(1)); return String.valueOf(Character.toChars(codePoint)); }); // Handle any remaining plain &#xxxx; entities result = result.replaceAll("&#(\\d+);", matcher -> { int codePoint = Integer.parseInt(matcher.group(1)); return String.valueOf(Character.toChars(codePoint)); }); System.out.println(result); } }
For Java 8 or older, use Matcher.appendReplacement:
import java.util.regex.Matcher; import java.util.regex.Pattern; public class EmojiConverter { public static void main(String[] args) { String input = "Last Search Results Bulletin Board Validations ⛔ 0 ⚡ 1 ⚠ 6 ? 0"; Pattern pattern = Pattern.compile("&#(\\d+);|&#(\\d+);"); Matcher matcher = pattern.matcher(input); StringBuffer sb = new StringBuffer(); while (matcher.find()) { // Grab the numeric group (either group 1 or 2) String numStr = matcher.group(1) != null ? matcher.group(1) : matcher.group(2); int codePoint = Integer.parseInt(numStr); matcher.appendReplacement(sb, String.valueOf(Character.toChars(codePoint))); } matcher.appendTail(sb); System.out.println(sb.toString()); } }
Solution 2: Remove the Unicode Entities Entirely
If you just want to delete these entities, use a regex that targets the exact entity formats.
Remove All Such Entities
This will delete any &#xxxx; or &#xxxx; pattern:
public class EntityRemover { public static void main(String[] args) { String input = "Last Search Results Bulletin Board Validations ⛔ 0 ⚡ 1 ⚠ 6 ? 0"; String cleaned = input.replaceAll("&#\\d+;", "") .replaceAll("&#\\d+;", ""); System.out.println(cleaned); // Output: Last Search Results Bulletin Board Validations 0 1 6 ? 0 } }
Remove Only Entities in Your Target Range (9728-9983)
If you want to delete only entities in the emoji range you specified, use a regex that matches that numeric range specifically:
public class EntityRemover { public static void main(String[] args) { String input = "Last Search Results Bulletin Board Validations ⛔ 0 ⚡ 1 ⚠ 6 ? 0"; // Regex to match 9728-9983: covers all values in your range String rangeRegex = "(972[8-9]|97[3-9]\\d|9[8-9]\\d{2}|998[0-3])"; String cleaned = input.replaceAll("&#" + rangeRegex + ";", "") .replaceAll("&#" + rangeRegex + ";", ""); System.out.println(cleaned); } }
Bonus: Prevent This Issue at the Source
To avoid having emojis converted to entities in the first place, make sure:
- Your database and table columns use the
utf8mb4character set (supports full Unicode, including 4-byte emojis). - Your web form doesn't automatically escape HTML entities for the text field (or unescape them before saving to the database).
内容的提问来源于stack exchange,提问作者nagaraju

