如何在Java中验证HTML代码片段?求支持闭合标签校验的类库
Hey there! I get exactly what you're looking for—you need to validate specific HTML snippets (like your example <div class=something> <iframesrc="link">) for unclosed tags, not just full websites or general syntax checks. Let's break down some practical Java-based solutions:
1. 用Jsoup(最常用的HTML解析库)
Jsoup is a go-to for HTML processing in Java, and it's perfect for snippet validation. While it automatically fixes malformed HTML, it also tracks parsing errors—including unclosed tags. Here's how to use it to catch those issues:
import org.jsoup.Jsoup; import org.jsoup.nodes.Document; import org.jsoup.parser.ParseError; import org.jsoup.parser.Parser; public class UnclosedTagChecker { public static void main(String[] args) { String yourSnippet = "<div class=something> <iframesrc=\"link\">"; // Set up parser to track errors Parser parser = Parser.htmlParser(); ParseErrorList errorList = ParseErrorList.tracking(10); // Track up to 10 issues // Parse the snippet Document doc = Jsoup.parse(yourSnippet, "", parser, errorList); // Check for unclosed tag errors if (!errorList.isEmpty()) { System.out.println("Found issues with your HTML snippet:"); for (ParseError error : errorList) { System.out.println("- " + error.getErrorMessage()); } } // Bonus: Compare original vs parsed/fixed HTML to see what was corrected System.out.println("\nOriginal snippet: " + yourSnippet); System.out.println("Parsed & fixed version: " + doc.html()); } }
When you run this, Jsoup will flag both the unclosed <div> and the malformed <iframesrc> (missing space between iframe and src) as parsing errors.
2. HTMLCleaner(另一个可靠的解析工具)
HTMLCleaner is another solid option that focuses on cleaning and validating HTML. It can report parsing errors directly, including unclosed tags:
import org.htmlcleaner.CleanerProperties; import org.htmlcleaner.HtmlCleaner; import org.htmlcleaner.TagNode; public class HtmlCleanerValidator { public static void main(String[] args) { String yourSnippet = "<div class=something> <iframesrc=\"link\">"; HtmlCleaner cleaner = new HtmlCleaner(); CleanerProperties props = cleaner.getProperties(); // Enable error reporting props.setReportErrors(true); props.setShowWarnings(true); // Parse the snippet TagNode root = cleaner.clean(yourSnippet); // Print out detected errors System.out.println("Validation issues found:"); for (Object error : cleaner.getErrors()) { System.out.println("- " + error.toString()); } } }
This will highlight unclosed elements and syntax mistakes just like Jsoup, with a slightly different error format.
3. 自定义栈实现(如果需要完全自定义逻辑)
If you want to build something lightweight without relying on external libraries, you can use a stack-based approach:
- Traverse each character in the HTML snippet
- Push opening tags (excluding self-closing ones like
<img>,<br>) onto the stack - When you hit a closing tag, pop the matching opening tag from the stack
- Any tags left in the stack at the end are unclosed
Note: This requires handling edge cases like nested tags, self-closing tags, and attribute values that might contain angle brackets—so it's more work than using a mature library. But it's an option if you need full control.
All these solutions work for individual HTML snippets, not just entire websites. The library-based approaches are the most reliable since they handle all the edge cases you might miss with a custom implementation.
内容的提问来源于stack exchange,提问作者Skolak

