You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java 8正则表达式可选组(含空格)多场景匹配需求问询

Java 8 Regex Pattern for Multi-Scenario Matching with Group Extraction

Hey there! Since you're new to regex and need a pattern that handles multiple text scenarios while capturing groups for extraction via matcher.find(), let's walk through building this step by step.

First, let's align on your example: the input "employeeName:*borgy*" should split into groups like (employeeName), (:), (*), (borgy), (*). We'll build a pattern that works for this case and other common variations too.

Step 1: Define the Regex Pattern

We'll create a pattern that accounts for these common scenarios:

  • Exact matches (e.g., "employeeId:12345")
  • Prefix wildcard matches (e.g., "department:*sales")
  • Suffix wildcard matches (e.g., "jobTitle:dev*")
  • Double wildcard matches (your example: "employeeName:*borgy*")

Here's the regex pattern with clear capture groups:

String regex = "^\\s*([a-zA-Z0-9_]+)\\s*(:)\\s*(\\*)?([^*]+)(\\*)?\\s*$";

Let's break down each group:

  • Group 1: ([a-zA-Z0-9_]+) → Captures the field name (supports letters, numbers, underscores; adjust the character class if your field names have other characters)
  • Group 2: (:) → Captures the colon separator
  • Group 3: (\\*)? → Optional capture for the prefix wildcard (matches zero or one *)
  • Group 4: ([^*]+) → Captures the actual value (anything except *)
  • Group 5: (\\*)? → Optional capture for the suffix wildcard (matches zero or one *)
  • The \\s* parts handle optional whitespace around the field name, colon, and value (remove these if you don't need to support whitespace)

Step 2: Java 8 Implementation Example

Here's how to use this pattern with Matcher.find() to extract each group:

import java.util.regex.Matcher;
import java.util.regex.Pattern;

public class RegexGroupExtractor {
    public static void main(String[] args) {
        // Test multiple scenarios
        String[] testInputs = {
            "employeeName:*borgy*",
            "employeeId:12345",
            "department:*sales",
            "jobTitle:dev*",
            "  location : *New York*  " // With whitespace
        };

        String regex = "^\\s*([a-zA-Z0-9_]+)\\s*(:)\\s*(\\*)?([^*]+)(\\*)?\\s*$";
        Pattern pattern = Pattern.compile(regex);

        for (String input : testInputs) {
            Matcher matcher = pattern.matcher(input);
            if (matcher.find()) {
                System.out.println("Input: " + input);
                System.out.println("Group 1 (Field Name): " + matcher.group(1));
                System.out.println("Group 2 (Separator): " + matcher.group(2));
                System.out.println("Group 3 (Prefix Wildcard): " + (matcher.group(3) != null ? matcher.group(3) : "None"));
                System.out.println("Group 4 (Value): " + matcher.group(4));
                System.out.println("Group 5 (Suffix Wildcard): " + (matcher.group(5) != null ? matcher.group(5) : "None"));
                System.out.println("---");
            } else {
                System.out.println("Input: " + input + " → No match found");
                System.out.println("---");
            }
        }
    }
}

Step 3: Customizing for Your Needs

If you need to support additional scenarios:

  • Different field name characters: Modify the first group's character class, e.g., ([a-zA-Z0-9_\\-]+) to allow hyphens.
  • Other wildcard symbols: Replace \\* with your target symbol (e.g., % for SQL-style wildcards).
  • Case-insensitive matching: Add Pattern.CASE_INSENSITIVE when compiling the pattern:
    Pattern pattern = Pattern.compile(regex, Pattern.CASE_INSENSITIVE);
    

Expected Output for Your Example

For the input "employeeName:*borgy*", the output will be:

Input: employeeName:borgy
Group 1 (Field Name): employeeName
Group 2 (Separator): :
Group 3 (Prefix Wildcard): *
Group 4 (Value): borgy
Group 5 (Suffix Wildcard): *

内容的提问来源于stack exchange,提问作者Borgy Manotoy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:20:59