PCRE后向断言中\G的工作机制?为何与Java Split结果有差异?
\G Differs from Java Great question—this comes down to core differences in how Java and PCRE handle the \G anchor and zero-length delimiters during split operations. Let's break this down clearly:
First, Java's Approach
In Java, the \G anchor is strictly tied to the end position of the last successful delimiter match, and it doesn't reset even when the search starting position shifts. Additionally, Java avoids infinite loops with zero-length delimiters by skipping consecutive zero-length matches, but \G stays anchored to the last match end.
For your regex (?<=\G..) and string abcdefghij:
- First match:
\Gstarts at the string's beginning (position 0). The assertion(?<=\G..)matches right afterab(position 2), splitting outab.\Gnow anchors to position 2. - Next match: We start searching at position 2.
\Gis still anchored to 2, so the assertion matches right aftercd(position 4), splitting outcd.\Gupdates to 4. - This repeats, splitting after every two characters, resulting in
["ab", "cd", "ef", "gh", "ij"].
PCRE's Approach (Delphi & regex101)
PCRE has two key behaviors that change the outcome:
- Zero-length delimiter handling: When a split matches a zero-length delimiter (like your positive lookbehind), PCRE automatically advances the next search start position by 1 character to avoid infinite loops.
\Ganchor reset: Unlike Java, PCRE's\Ganchor binds to the current search start position, not the end of the last match.
Let's walk through your example step by step:
- First split:
- Start at position 0.
\Ganchors here, so(?<=\G..)matches position 2 (afterab). Split outab. - Since this is a zero-length match, the next search starts at position 3, and
\Gresets to anchor at 3.
- Start at position 0.
- Second split:
- Start at position 3.
\Ganchors here, so(?<=\G..)matches position 3+2=5 (afterde). Split outcde(from position 2 to 5). - Next search starts at position 6,
\Gresets to 6.
- Start at position 3.
- Third split:
- Start at position 6.
\Ganchors here, so(?<=\G..)matches position 6+2=8 (aftergh). Split outfgh(from position 5 to 8). - Next search starts at position 9,
\Gresets to 9.
- Start at position 6.
- Final segment:
- Starting at position 9,
(?<=\G..)would need to match position 11 (beyond the string's end), so no more splits. The remainingijis added as the last element.
- Starting at position 9,
This gives you the result you saw: ["ab", "cde", "fgh", "ij"].
Key Takeaway
The gap between Java and PCRE here boils down to how each engine defines the \G anchor in split contexts: Java keeps \G tied to the last match's end, while PCRE resets \G to the current search start (especially after zero-length matches). This changes where the split points land entirely.
内容的提问来源于stack exchange,提问作者Sebastian Proske

