You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java正则能否用捕获组值定义其他捕获组?场景验证

问题描述

我有一个符合以下格式的字符串:

"<N>|<M>|<word_of_size_N><word_of_size_M>"

是否存在单个正则表达式,能够找到<N>和<M>的值,并利用这些值提取<word_of_size_N>和<word_of_size_M>?

我尝试了以下Java代码:

import java.util.regex.*;

public class RegexExample {
    public static void main(String[] args) {
        String input = "5|3|hello123";
        String regex = "^(\\d+)\\|(\\d+)\\|(\\w{\\1})(\\w{\\2})$";
        Pattern pattern = Pattern.compile(regex);
        Matcher matcher = pattern.matcher(input);
        
        if (matcher.matches()) {
            String n = matcher.group(1);
            String m = matcher.group(2);
            String wordN = matcher.group(3);
            String wordM = matcher.group(4);
            System.out.println("The input string matches the regex pattern.");
            System.out.println("N = " + n + ", M = " + m);
            System.out.println("word of size N = " + wordN);
            System.out.println("word of size M = " + wordM);
        } else {
            System.out.println("The input string does not match the regex pattern.");
        }
    }
}

运行后抛出异常:

Exception in thread "main" java.util.regex.PatternSyntaxException: Illegal repetition near index 17
^(\d+)\|(\d+)\|(\w{\1})(\w{\2})$
                 ^
at java.base/java.util.regex.Pattern.error(Pattern.java:2028)
at java.base/java.util.regex.Pattern.closure(Pattern.java:3323)
    at java.base/java.util.regex.Pattern.sequence(Pattern.java:2214)
    at java.base/java.util.regex.Pattern.expr(Pattern.java:2069)
    at java.base/java.util.regex.Pattern.group0(Pattern.java:3060)
    at java.base/java.util.regex.Pattern.sequence(Pattern.java:2124)
    at java.base/java.util.regex.Pattern.expr(Pattern.java:2069)
    at java.base/java.util.regex.Pattern.compile(Pattern.java:1783)
    at java.base/java.util.regex.Pattern.<init>(Pattern.java:1429)
    at java.base/java.util.regex.Pattern.compile(Pattern.java:1069)
    at RegexExample.main(RegexExample.java:7)

我知道两步法可以实现:

  • 用正则或字符串操作提取<N>和<M>的数值
  • 根据数值截取对应长度的子串

但我好奇正则是否支持这种自引用方式,另外想知道有没有其他正则引擎支持该特性?

解答

Java正则的局限性

你写的(\w{\1})这种语法在Java里不成立,因为Java的正则引擎属于传统NFA引擎,不支持在量词中引用前面捕获组的内容。引擎会把{后面的\1当成普通字符,但量词要求{后必须是固定数字或数字范围,所以直接触发语法错误。

Java中能否用单个正则实现?

不行,必须拆分两步:

  1. 先用正则^(\\d+)\\|(\\d+)\\|(\\w+)$匹配整个字符串,提取出N(group(1))和M(group(2))的数值;
  2. 再从第三个捕获组(group(3))中,截取前N位作为<word_of_size_N>,后M位作为<word_of_size_M>。

支持该特性的正则引擎

有不少增强型正则引擎支持动态量词(即用量词引用前面的捕获组),比如:

  • PCRE系列:Perl、PHP、Rust的regex库等都基于PCRE,直接支持^\d+\|\d+\|\w{\1}\w{\2}$这类写法,能一次性匹配并提取目标内容。
  • .NET正则引擎:C#等.NET语言的正则同样支持这种动态引用,写法和PCRE一致。
  • Python第三方regex库:Python标准库的re不支持,但第三方的regex库兼容PCRE特性,可实现该需求。

内容的提问来源于stack exchange,提问作者Damiano Caprari

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 18:50:38