Python中如何捕获不在非捕获组内的目标模式?
Hey there! Let's figure out how to capture those target tokens while skipping the common contractions you want to avoid.
You already have a regex to match the contractions you don't want ('ll, 've, 't), and now you need the opposite: grab tokens like 'v and 'w that aren't part of those excluded contractions. The trick here is using a negative lookahead to rule out the unwanted patterns before matching what you do want.
Here's the working code example:
import re x = "I'll be going home I've the 'v ' isn't want I want to split but I want to catch tokens like 'v and 'w ' ." # Breakdown of the regex pattern: # (?i) = case-insensitive matching (so 'V' or 'W' would also be caught) # ' = matches the single quote # (?!ve|ll|t) = negative lookahead: ensures the characters after ' aren't "ve", "ll", or "t" # \w = matches a single alphanumeric character (perfect for your 'v/'w use case) pattern = r"(?i)'(?!ve|ll|t)\w" matches = re.findall(pattern, x) print(matches) # Output: ["'v", "'v", "'w"]
Let me break down the regex piece by piece to make it clear:
(?i): Turns on case-insensitive matching, so you don't have to worry about uppercase vs lowercase letters.': Literally matches the single quote at the start of your target tokens.(?!ve|ll|t): This is the key part—it checks that the characters immediately after the single quote are NOT "ve", "ll", or "t". This skips exactly the contractions you want to avoid.\w: Matches a single letter (or number, though your use case is letters) to capture the token you care about. If you ever need to capture longer tokens (like'xyz), just change this to\w+.
If you want to make sure you're only capturing standalone tokens (not part of a longer word), you can add a word boundary at the end: (?i)'(?!ve|ll|t)\w\b. This prevents matches like 'v from 'vw (though that scenario doesn't come up in your sample text).
内容的提问来源于stack exchange,提问作者alvas

