正则表达式匹配WordA与WordB间数字的问题求助
The issue with your original regex (?<=WordA ).*(?= WordB) is that the .* is a greedy quantifier—it matches as much text as possible, which means it grabs everything from the first WordA all the way to the last WordB, including the irrelevant middle content. Here are two reliable solutions to get exactly the number sets you want:
Solution 1: Use a Non-Greedy Quantifier
Switch to the non-greedy version of the wildcard (.*?), which matches as little text as needed to reach the next WordB. This is the simplest fix for your specific input:
(?<=WordA ).*?(?= WordB)
How it works:
(?<=WordA ): Positive lookbehind to ensure we start right afterWordA.*?: Non-greedy wildcard that stops at the first occurrence ofWordB(?= WordB): Positive lookahead to ensure we end right beforeWordB
When run with global matching (to find all occurrences), this regex will capture:
1 2 3 4(between the firstWordAand firstWordB)13 14 15 16(between the secondWordAand finalWordB)
Solution 2: Robust Match with Negative Lookahead
If you want to guarantee no other WordA or WordB appear in the captured content (for edge cases where those words might pop up between your target pairs), use a negative lookahead to exclude those sequences:
(?<=WordA )(?:(?!WordA|WordB).)*(?= WordB)
How it works:
(?:(?!WordA|WordB).)*: Matches any character only if it doesn’t start aWordAorWordBsequence, repeating until we hitWordB- This prevents accidental inclusion of unrelated
WordA/WordBinstances in your captured numbers
Example Usage (Python)
Here’s how you’d implement the non-greedy regex in Python to extract your target sets:
import re input_text = "WordA 1 2 3 4 WordB 5 6 7 8 WordC 9 10 11 12 WordA 13 14 15 16 WordB" matches = re.findall(r'(?<=WordA ).*?(?= WordB)', input_text) print(matches) # Output: ['1 2 3 4', '13 14 15 16']
内容的提问来源于stack exchange,提问作者abracadab

