You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则表达式提取路径字符串的所有层级前缀匹配项?

Extracting Hierarchical Prefixes with Python Regex

Your current regex isn't working because the non-greedy .+? matches the shortest possible substring ending with Node[X], which causes subsequent matches to start from the end of the previous match instead of the beginning of the string. This results in partial paths rather than cumulative prefixes.

Correct Regex Solution

To capture all cumulative hierarchical prefixes, we can use a positive lookahead that anchors to the start of the string and matches each full prefix incrementally. Here's the working code:

import re

string = 'TreeModel/Node/Node[1]/Node[4]/Node[1]'
# Regex pattern to match all cumulative prefixes from the start
pattern = r'(?=(^TreeModel/Node(?:/Node\[[1-9]\])*))'

# Get all matches (will include duplicates)
matches = re.findall(pattern, string)
# Remove duplicates and sort by length to get the correct order
unique_prefixes = sorted(list(set(matches)), key=len)

print(unique_prefixes)

Output:

['TreeModel/Node', 'TreeModel/Node/Node[1]', 'TreeModel/Node/Node[1]/Node[4]', 'TreeModel/Node/Node[1]/Node[4]/Node[1]']

How This Works

  • (?=...): Positive lookahead assertion that checks for a match without consuming characters, allowing us to capture multiple prefixes starting from the string's beginning.
  • ^TreeModel/Node: Anchors the match to the start of the string and matches the initial base path.
  • (?:/Node\[[1-9]\])*: Non-capturing group that matches zero or more subsequent /Node[X] segments (where X is a digit from 1-9). The * allows us to match all incremental prefixes.
  • We use set() to remove duplicate matches (since the lookahead will trigger at multiple character positions for the same prefix) and sorted() with key=len to ensure the prefixes are ordered from shortest to longest.

Alternative Non-Regex Approach (More Readable)

If regex feels overcomplicated, splitting the string into segments and building prefixes manually is often clearer:

string = 'TreeModel/Node/Node[1]/Node[4]/Node[1]'
parts = string.split('/')
prefixes = ['/'.join(parts[:i+1]) for i in range(len(parts))]

print(prefixes)

This produces the same output and is easier to understand at a glance.

内容的提问来源于stack exchange,提问作者deltascience

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:44:34