You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PCRE后向断言中\G的工作机制?为何与Java Split结果有差异?

Why PCRE's Split Behavior with \G Differs from Java

Great question—this comes down to core differences in how Java and PCRE handle the \G anchor and zero-length delimiters during split operations. Let's break this down clearly:

First, Java's Approach

In Java, the \G anchor is strictly tied to the end position of the last successful delimiter match, and it doesn't reset even when the search starting position shifts. Additionally, Java avoids infinite loops with zero-length delimiters by skipping consecutive zero-length matches, but \G stays anchored to the last match end.

For your regex (?<=\G..) and string abcdefghij:

  • First match: \G starts at the string's beginning (position 0). The assertion (?<=\G..) matches right after ab (position 2), splitting out ab. \G now anchors to position 2.
  • Next match: We start searching at position 2. \G is still anchored to 2, so the assertion matches right after cd (position 4), splitting out cd. \G updates to 4.
  • This repeats, splitting after every two characters, resulting in ["ab", "cd", "ef", "gh", "ij"].

PCRE's Approach (Delphi & regex101)

PCRE has two key behaviors that change the outcome:

  1. Zero-length delimiter handling: When a split matches a zero-length delimiter (like your positive lookbehind), PCRE automatically advances the next search start position by 1 character to avoid infinite loops.
  2. \G anchor reset: Unlike Java, PCRE's \G anchor binds to the current search start position, not the end of the last match.

Let's walk through your example step by step:

  1. First split:
    • Start at position 0. \G anchors here, so (?<=\G..) matches position 2 (after ab). Split out ab.
    • Since this is a zero-length match, the next search starts at position 3, and \G resets to anchor at 3.
  2. Second split:
    • Start at position 3. \G anchors here, so (?<=\G..) matches position 3+2=5 (after de). Split out cde (from position 2 to 5).
    • Next search starts at position 6, \G resets to 6.
  3. Third split:
    • Start at position 6. \G anchors here, so (?<=\G..) matches position 6+2=8 (after gh). Split out fgh (from position 5 to 8).
    • Next search starts at position 9, \G resets to 9.
  4. Final segment:
    • Starting at position 9, (?<=\G..) would need to match position 11 (beyond the string's end), so no more splits. The remaining ij is added as the last element.

This gives you the result you saw: ["ab", "cde", "fgh", "ij"].

Key Takeaway

The gap between Java and PCRE here boils down to how each engine defines the \G anchor in split contexts: Java keeps \G tied to the last match's end, while PCRE resets \G to the current search start (especially after zero-length matches). This changes where the split points land entirely.

内容的提问来源于stack exchange,提问作者Sebastian Proske

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 04:14:10