使用正则表达式掩码指定XML元素敏感信息的技术问询
Fixing Your Regex for Sensitive Data Masking in XML
Let's break down your problem and fix the regex step by step—your current pattern won't work as intended, so we'll adjust it to meet your requirements properly.
What's Wrong With Your Current Regex?
Your existing expression:
final String expressionTemplate = "(?<=<element name=\"(?i:XYZ)\"[a-z?]>)(.*?)(?=</element>)";
has a critical flaw in the lookbehind part:
[a-z?]>only matches a single lowercase letter or a question mark followed by a closing angle bracket. This can't match valid attributes liketype="string"that come after thenameattribute. It will completely fail to target the content inside elements with proper additional attributes.- Additionally, Java 8 and earlier don't support variable-length lookbehinds, so even if you fixed the attribute matching, this pattern would throw an error in older Java versions.
Correct Regex & Implementation
Instead of relying on lookbehinds, use capturing groups to preserve the parts of the XML you want to keep, while replacing only the sensitive content. Here's the solution:
public class XmlMasking { public static void main(String[] args) { String input = "<element name=\"XYZ\" type=\"string\">9032905776</element>"; // Regex breakdown: // 1. Capture the opening tag with name="XYZ" (case-insensitive) and any other attributes // 2. Capture the sensitive content we want to replace // 3. Capture the closing tag String regex = "(<element name=\"(?i:XYZ)\"[^>]*>)(.*?)(</element>)"; // Replace with the opening tag, *** for content, and closing tag String maskedOutput = input.replaceAll(regex, "$1***$3"); System.out.println(maskedOutput); // Output: <element name="XYZ" type="string">***</element> } }
Key Details About the Regex:
(?i:XYZ): Makes theXYZmatch case-insensitive (soxyz,XyZ, etc., will all trigger masking)[^>]*: Matches any characters except>(this safely captures all additional attributes liketype="string"without breaking the tag structure)- Capturing groups (
$1,$3): Preserve the opening and closing tags exactly as they were, so we don't modify any attributes—only the element content.
Edge Cases It Handles:
- Elements with multiple attributes:
<element name="XYZ" type="string" required="true">12345</element>becomes<element name="XYZ" type="string" required="true">***</element> - Case-variant
namevalues:<element name="xyz" type="int">6789</element>is correctly masked - Non-target elements:
<element name="ABC" type="string">1111</element>remains unchanged
内容的提问来源于stack exchange,提问作者Namburi Manikanta
相关产品推荐
相关产品推荐

