Java中如何捕获带属性的<s:message>标签并按code值替换内容?
Got it, so you need to scan an HTML string or file for those <s:message> tags and swap them out based on their code attribute value. Here's how you can do it effectively in Java, with two common approaches: regex for simple cases, and an HTML parser for more robust scenarios.
1. Quick Solution: Regular Expressions
If your HTML is well-formed and the <s:message> tags follow a consistent structure (like your example), regex works great. We'll create a pattern that matches the tag, captures the code value, and then replaces each match dynamically.
First, the regex pattern to target the tags:
String tagPattern = "<s:message\\s+code=\"([^\"]+)\"(?:\\s+arguments=\"[^\"]+\")?\\s*/>";
Breakdown of the pattern:
<s:message\\s+: Matches the opening of the tag with any whitespace after it.code=\"([^\"]+)\": Captures thecodeattribute's value (everything inside the quotes).(?:\\s+arguments=\"[^\"]+\")?: Optional part to account for theargumentsattribute (the?:makes it a non-capturing group since we might not need that value).\\s*/>: Matches the self-closing end of the tag, allowing for any trailing whitespace.
Full Regex Implementation
Here's a complete method to handle string input:
import java.util.regex.Matcher; import java.util.regex.Pattern; public class MessageTagProcessor { public static void main(String[] args) { String input = "Some text........... <s:message code=\"code1\" arguments=\"${arg1,arg2}\" />.. some text ........ some text ....... <s:message code=\"code2\" />..........."; Pattern pattern = Pattern.compile("<s:message\\s+code=\"([^\"]+)\"(?:\\s+arguments=\"[^\"]+\")?\\s*/>"); Matcher matcher = pattern.matcher(input); StringBuffer result = new StringBuffer(); while (matcher.find()) { String code = matcher.group(1); // Map code to your replacement text String replacement = switch(code) { case "code1" -> "test1"; case "code2" -> "test2"; default -> ""; // Or keep the original tag if code is unrecognized }; // Use quoteReplacement to avoid issues with special characters in replacement matcher.appendReplacement(result, Matcher.quoteReplacement(replacement)); } matcher.appendTail(result); System.out.println(result.toString()); } }
2. Robust Solution: Use Jsoup HTML Parser
Regex can break if your HTML has malformed tags, varying attribute orders, or nested elements. For production-grade code, use Jsoup—it's a powerful HTML parser that handles all these edge cases.
First, add Jsoup to your project (via Maven, Gradle, or download the JAR). Then:
import org.jsoup.Jsoup; import org.jsoup.nodes.Document; import org.jsoup.nodes.Element; import org.jsoup.select.Elements; public class JsoupTagReplacer { public static void main(String[] args) throws Exception { String inputHtml = "Your HTML content here (or read from file)"; // Parse the HTML Document doc = Jsoup.parse(inputHtml); // Find all <s:message> tags Elements messageTags = doc.select("s|message"); for (Element tag : messageTags) { String code = tag.attr("code"); String replacement = getReplacement(code); // Replace the tag with your text tag.replaceWith(Jsoup.parse(replacement).textNode()); } // Get the modified HTML String outputHtml = doc.html(); System.out.println(outputHtml); } private static String getReplacement(String code) { return switch(code) { case "code1" -> "test1"; case "code2" -> "test2"; default -> tag.outerHtml(); // Fallback to original tag if code is unknown }; } }
Handling File Input
If your HTML is stored in a file, read it into a string first using Java NIO:
import java.nio.file.Files; import java.nio.file.Paths; String fileContent = new String(Files.readAllBytes(Paths.get("path/to/your/html/file.html")));
Key Notes
- Regex vs Jsoup: Use regex only if you're certain the HTML is simple and consistent. Jsoup is better for real-world HTML where tags might be messy.
- Arguments Attribute: If you need to use the
argumentsvalue in your replacement, adjust the regex to capture it (add another group likearguments=\"([^\"]+)\") or usetag.attr("arguments")in Jsoup. - Special Characters: Always use
Matcher.quoteReplacement()when using regex to avoid issues with characters like$in your replacement text.
内容的提问来源于stack exchange,提问作者emilly

