如何区分<、>是运算符还是HTML标签括号?求替换运算符的Java代码
Great question! When working with HTML, it’s easy to mix up angle brackets used for HTML tag syntax (like <html>) and those used as comparison operators (like < 30). The safest way to handle this is to use a dedicated HTML parser instead of regex—regex can’t reliably parse HTML structure, leading to bugs with nested tags, entities, or comments.
Here’s a step-by-step solution using Jsoup (a popular Java HTML parser):
1. Core Concept
HTML tags are part of the document’s structural markup, while operator brackets live inside text content. Our goal is to:
- Parse the HTML to separate structural elements from user-facing text nodes.
- Modify only the text nodes, leaving tag brackets completely untouched.
2. Add Jsoup Dependency
First, include Jsoup in your project. For Maven, add this to your pom.xml:
<dependency> <groupId>org.jsoup</groupId> <artifactId>jsoup</artifactId> <version>1.17.2</version> </dependency>
3. Java Code to Replace Operators
This code parses the HTML, traverses only text nodes, and replaces < with "lt" and > with "gt":
import org.jsoup.Jsoup; import org.jsoup.nodes.Document; import org.jsoup.nodes.Node; import org.jsoup.nodes.TextNode; public class HtmlOperatorReplacer { public static void main(String[] args) { String htmlContent = "<html xmlns=\"http://www.w3.org/1999/xhtml\" xmlns:epub=\"http://www.idpf.org/2007/ops\" lang=\"en\" xmlns:ns=\"http://www.w3.org/2001/10/synthesis\" xmlns:m=\"http://www.w3.org/1998/Math/MathML\" xml:lang=\"en\"> <head> </head> <body> <div id=\"1234\">My name is ABC and my age is < 30. </div> </body> </html>"; // Parse the HTML into a structured DOM tree Document doc = Jsoup.parse(htmlContent); // Traverse and modify only text nodes processTextNodes(doc); // Output the updated HTML System.out.println(doc.html()); } private static void processTextNodes(Node node) { // Target only text nodes (these hold the user-readable content) if (node instanceof TextNode) { TextNode textNode = (TextNode) node; String originalText = textNode.text(); // Replace operator brackets with the requested strings // Jsoup automatically unescapes entities like < to '<' in text() String updatedText = originalText.replace("<", "lt").replace(">", "gt"); // Update the text node with modified content textNode.text(updatedText); } // Recursively check all child nodes to cover nested content for (Node child : node.childNodes()) { processTextNodes(child); } } }
4. Key Details
- Structural Safety: By using Jsoup’s DOM parsing, we never touch HTML tags—only the text content inside elements.
- Entity Handling: Jsoup automatically converts HTML entities (like
<) to their actual character (<) in thetext()method, so we don’t have to handle entity codes manually. - Nested Content Support: The recursive traversal ensures we modify text even inside nested elements (like the
<div>in your example).
5. Why Regex Is a Bad Idea
Regex can’t reliably tell the difference between a < in a tag (e.g., <div>) and a < in user text (e.g., < 30). It might accidentally modify tag syntax, breaking the entire HTML structure. A dedicated parser like Jsoup eliminates this risk entirely.
内容的提问来源于stack exchange,提问作者Vikas Deshmukh

