如何用PyParsing从指定字符串中提取所有符合规则的重叠匹配项?
Your original code only returns the first match because parseString expects the entire input to conform to your rule. Since your string has non-matching characters (like trailing '0's and sequences of '1's without '0's in between), the parser stops at the first non-match. Additionally, your expected results include overlapping matches (like '10001' and '1001' sharing the middle '1'), which standard left-to-right parsing won't catch by default.
Here are two PyParsing-based solutions to get your expected output:
Solution 1: Using Regex with Positive Lookahead (Efficient for Overlapping Matches)
PyParsing's Regex class lets you leverage regex's lookahead feature to find all overlapping matches without consuming characters, which is perfect for your case:
from pyparsing import Regex s = '10001001110100000' # Regex pattern with positive lookahead to capture overlapping matches match_pattern = Regex(r'(?=(10+1))') matches = [] # Iterate through all matches found by scanString for tokens, _, _ in match_pattern.scanString(s): matches.append(tokens[1]) print(matches) # Output: ['10001', '1001', '101']
The regex (?=(10+1)) uses a positive lookahead to identify all sequences of 1 followed by one or more 0s followed by another 1, without consuming the characters—allowing overlapping matches to be detected.
Solution 2: Pure PyParsing with Iterative Scanning (No Regex)
If you prefer to avoid regex, you can iterate through each position in the string and attempt to parse starting at that index. This brute-force approach works for your use case:
from pyparsing import Literal, OneOrMore s = '10001001110100000' one = Literal('1') zero = Literal('0') # Define the pattern: 1 -> one or more 0s -> 1 expr = one + OneOrMore(zero) + one matches = [] for i in range(len(s)): try: # Try to parse starting at index i, don't require full string match result = expr.parseString(s[i:], parseAll=False) if result: matches.append(''.join(result)) except: # Skip positions where the pattern doesn't match pass print(matches) # Output: ['10001', '1001', '101']
This approach checks every possible starting position, so it captures overlapping matches and skips non-matching sections naturally.
Why Your Original Code Failed
parseStringexpects the entire input to match your rule (ZeroOrMore(Group(expr))). Since your string has parts that don't fitexpr(like leading/trailing '0's or consecutive '1's), the parser stops after the first valid match.- Even if you used
scanStringwith your originalexpr, it wouldn't find overlapping matches because it consumes characters as it parses, so it would miss the '1001' match that starts at the end of the first match.
内容的提问来源于stack exchange,提问作者Raphael

