如何高效去除Java HashMap值中包含的HTML标签?
实现方案分两类,可根据业务场景选择:
方案1:正则表达式实现(无第三方依赖,性能最高,适合简单HTML标签场景)
针对格式规范的简单HTML标签,直接预编译正则匹配所有<xxx>格式的标签替换为空即可,性能优于HTML解析类方案:
import java.util.HashMap; import java.util.Map; import java.util.regex.Pattern; public class RemoveHtmlTag { // 预编译正则,避免每次调用重复编译,大幅提升性能 private static final Pattern HTML_TAG_PATTERN = Pattern.compile("<[^>]+>"); public static void main(String[] args) { HashMap<String, String> hm = new HashMap<>(); hm.put("A", "Apple"); hm.put("B", "<b>Ball</b>"); hm.put("C", "Cat"); hm.put("D", "Dog"); hm.put("E", "<h1>Elephant</h1>"); // 遍历HashMap处理所有值 for (Map.Entry<String, String> entry : hm.entrySet()) { String originalValue = entry.getValue(); if (originalValue != null) { String cleanedValue = HTML_TAG_PATTERN.matcher(originalValue).replaceAll(""); entry.setValue(cleanedValue); } } // 打印验证结果 hm.forEach((k,v) -> System.out.println(k + " : " + v)); } }
注意:该方案仅适用于标签格式规范、</>不会出现在标签属性或文本内容中的场景,没有额外的DOM解析开销,性能最优。
方案2:Jsoup解析实现(容错性强,适合复杂HTML场景)
如果值中可能存在带属性的标签、嵌套标签、不规范HTML写法,用Jsoup提取纯文本的方案更稳妥,不会出现正则匹配失误的问题:
import org.jsoup.Jsoup; import java.util.HashMap; import java.util.Map; public class RemoveHtmlTag { public static void main(String[] args) { HashMap<String, String> hm = new HashMap<>(); hm.put("A", "Apple"); hm.put("B", "<b class='text'>Ball</b>"); hm.put("C", "Cat"); hm.put("D", "Dog"); hm.put("E", "<h1>Elephant<span>Test</span></h1>"); // 遍历处理所有值 for (Map.Entry<String, String> entry : hm.entrySet()) { String originalValue = entry.getValue(); if (originalValue != null) { String cleanedValue = Jsoup.parse(originalValue).text(); entry.setValue(cleanedValue); } } // 打印验证结果 hm.forEach((k,v) -> System.out.println(k + " : " + v)); } }
注意:该方案需要引入Jsoup依赖,性能略低于正则方案,但容错性极强,不会出现标签清理不干净的问题。
内容的提问来源于stack exchange,提问作者Sadina Khatun
相关产品推荐
相关产品推荐

