如何用Python正则表达式提取路径字符串的所有层级前缀匹配项?
Your current regex isn't working because the non-greedy .+? matches the shortest possible substring ending with Node[X], which causes subsequent matches to start from the end of the previous match instead of the beginning of the string. This results in partial paths rather than cumulative prefixes.
Correct Regex Solution
To capture all cumulative hierarchical prefixes, we can use a positive lookahead that anchors to the start of the string and matches each full prefix incrementally. Here's the working code:
import re string = 'TreeModel/Node/Node[1]/Node[4]/Node[1]' # Regex pattern to match all cumulative prefixes from the start pattern = r'(?=(^TreeModel/Node(?:/Node\[[1-9]\])*))' # Get all matches (will include duplicates) matches = re.findall(pattern, string) # Remove duplicates and sort by length to get the correct order unique_prefixes = sorted(list(set(matches)), key=len) print(unique_prefixes)
Output:
['TreeModel/Node', 'TreeModel/Node/Node[1]', 'TreeModel/Node/Node[1]/Node[4]', 'TreeModel/Node/Node[1]/Node[4]/Node[1]']
How This Works
(?=...): Positive lookahead assertion that checks for a match without consuming characters, allowing us to capture multiple prefixes starting from the string's beginning.^TreeModel/Node: Anchors the match to the start of the string and matches the initial base path.(?:/Node\[[1-9]\])*: Non-capturing group that matches zero or more subsequent/Node[X]segments (where X is a digit from 1-9). The*allows us to match all incremental prefixes.- We use
set()to remove duplicate matches (since the lookahead will trigger at multiple character positions for the same prefix) andsorted()withkey=lento ensure the prefixes are ordered from shortest to longest.
Alternative Non-Regex Approach (More Readable)
If regex feels overcomplicated, splitting the string into segments and building prefixes manually is often clearer:
string = 'TreeModel/Node/Node[1]/Node[4]/Node[1]' parts = string.split('/') prefixes = ['/'.join(parts[:i+1]) for i in range(len(parts))] print(prefixes)
This produces the same output and is easier to understand at a glance.
内容的提问来源于stack exchange,提问作者deltascience

