You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Groovy正则问题:String.split()正则与Matcher正则结果不一致

Extract Content Between Markers Using String.split() (No Matcher Allowed)

Got it, let's work through this since you can't use a Matcher and need to rely on String.split() to pull all content between your start/end markers into a list. Here's a straightforward, reliable approach:

Step-by-Step Logic

The core idea is to split twice: first on your start marker to isolate segments that follow each occurrence of the start, then split each of those segments on the end marker to grab the content in between.

Example Code (Java)

Let's use your sample input structure to make this concrete. Assume your start marker is something like START and end marker is END (adjust these to match your actual markers):

import java.util.ArrayList;
import java.util.List;
import java.util.regex.Pattern;

public class MarkerExtractor {
    public static void main(String[] args) {
        String input = "random text START A:STUFF1 B:MORE2 C:THAT3 END other junk START A:STUFF4 B:MORE5 C:THAT6 END more text START A:STUFF7 B:MORE8 C:THAT9 END";
        
        // Define your markers (use Pattern.quote() to escape special regex characters)
        String startMarker = "START";
        String endMarker = "END";
        
        // First split: split on the start marker to get segments after each start
        String[] postStartSegments = input.split(Pattern.quote(startMarker));
        
        List<String> extractedContent = new ArrayList<>();
        
        // Skip the first segment (it's everything before the first start marker)
        for (int i = 1; i < postStartSegments.length; i++) {
            // Split each post-start segment on the end marker, only split once
            String[] splitOnEnd = postStartSegments[i].split(Pattern.quote(endMarker), 2);
            
            // If we found a valid end marker, add the content between start/end
            if (splitOnEnd.length > 0) {
                // Trim whitespace if needed (adjust based on your input's formatting)
                extractedContent.add(splitOnEnd[0].trim());
            }
        }
        
        // Result will be your list of matched content:
        // [A:STUFF1 B:MORE2 C:THAT3, A:STUFF4 B:MORE5 C:THAT6, A:STUFF7 B:MORE8 C:THAT9]
        System.out.println(extractedContent);
    }
}

Key Details to Keep in Mind

  • Pattern.quote(): This is non-negotiable if your markers contain regex-special characters (like [, ], ., or *). It escapes them so they're treated as literal text instead of regex syntax.
  • Split Limit (2): Using 2 as the second parameter when splitting on the end marker ensures we only split once. This prevents issues if your content accidentally contains the end marker (though you should avoid that scenario if possible!).
  • Edge Case Handling:
    • If there are no start markers in the input, the result list stays empty.
    • If a start marker has no corresponding end marker, the code will add everything from that start marker to the end of the input. You can add a check to skip these (e.g., only add if splitOnEnd.length == 2) if that's not desired.

Quick Customization Tips

  • Swap startMarker and endMarker with your actual tags (e.g., <!--BEGIN--> and <!--END-->).
  • Remove the .trim() call if you need to preserve leading/trailing whitespace between the markers.

内容的提问来源于stack exchange,提问作者rboy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:33:34