Java中如何按指定字符数截取HTML字符串并保留原有格式避免单词中断
实现方案
核心思路
不要自己用正则匹配解析HTML,容错率极低,建议使用成熟的HTML解析库Jsoup完成需求,既能保留原有标签、属性结构,又能精准定位纯文本的截断位置。
步骤1:引入Jsoup依赖
如果用Maven管理项目,在pom.xml中添加以下依赖:
<dependency> <groupId>org.jsoup</groupId> <artifactId>jsoup</artifactId> <version>1.17.2</version> <!-- 可替换为最新稳定版 --> </dependency>
步骤2:具体实现逻辑
- 先将输入的HTML字符串解析为Jsoup的DOM树对象,所有原有标签、属性都会完整保留
- 遍历DOM的文本节点,累计纯文本的长度,直到累计长度达到设定的n值
- 触发长度阈值后,向前回溯到最近的空白分隔符(空格、制表符、换行等),避免截断在单词内部
- 截断超出的文本内容后,Jsoup会自动补全所有未闭合的标签,保证输出HTML结构合法
示例代码
import org.jsoup.Jsoup; import org.jsoup.nodes.Document; import org.jsoup.nodes.Node; import org.jsoup.nodes.TextNode; public class HtmlTruncator { public static String truncateHtml(String html, int maxLength) { Document doc = Jsoup.parse(html); doc.outputSettings().prettyPrint(false); // 保留原有格式不自动美化 int[] currentLength = {0}; boolean[] isTruncated = {false}; // 递归遍历DOM节点处理 traverseNode(doc.body(), currentLength, maxLength, isTruncated); return doc.body().html(); } private static void traverseNode(Node node, int[] currentLength, int maxLength, boolean[] isTruncated) { if (isTruncated[0]) { node.remove(); return; } if (node instanceof TextNode) { TextNode textNode = (TextNode) node; String text = textNode.text(); int remaining = maxLength - currentLength[0]; if (text.length() <= remaining) { currentLength[0] += text.length(); return; } // 找到最近的空白位置截断,避免拆分单词 int cutIndex = text.lastIndexOf(' ', remaining); if (cutIndex == -1) { cutIndex = remaining; // 如果剩余部分没有空格直接截断 } textNode.text(text.substring(0, cutIndex) + " "); currentLength[0] += cutIndex; isTruncated[0] = true; return; } // 遍历子节点 for (int i = 0; i < node.childNodes().size(); ) { Node child = node.childNode(i); traverseNode(child, currentLength, maxLength, isTruncated); if (!isTruncated[0]) { i++; } else { // 截断后删除后面所有子节点 for (int j = i + 1; j < node.childNodes().size(); j++) { node.childNode(j).remove(); } break; } } } public static void main(String[] args) { String inputHtml = "<p><strong id=\"faq_q1\">Q. What is the basic difference between Bachelor of Technology and Bachelor of Engineering?</strong></p>"; int n = 20; String result = truncateHtml(inputHtml, n); System.out.println(result); // 输出结果和题目示例完全一致:<p><strong id="faq_q1">Q. What is the basic </strong></p> } }
注意事项
- 如果需要处理中文、中日韩等无空格分隔的文本,可以调整截断逻辑,改为按语义或者直接按字符数截断即可
- Jsoup的parse方法默认会补全缺失的HTML根标签,如果需要完全保留输入的原始结构,可以使用
Parser.parseFragment()方法解析HTML片段
内容的提问来源于stack exchange,提问作者xerxes01
相关产品推荐
相关产品推荐

