Python正则:如何修改现有规则排除含'da cui'的匹配行?
Hey Anna, let's work through this regex problem step by step!
First off, you were on the right track with that negative lookahead pattern ^((?!da\scui).)*$ — that's exactly how we exclude strings containing a specific substring in regex. The trick is integrating it properly with your existing pattern instead of using it alone.
Let's break down your original regex
Your current pattern \]\s1\s([\w\s]+)\s2\s([\w\s]+) does a good job matching the ] 1 words 2 words structure, but the [\w\s]+ captures will happily include text with da cui since there's no restriction on that.
Two solid solutions to exclude da cui
Option 1: Exclude the entire string if it contains da cui
This is the simplest approach if you want to reject any line that has da cui anywhere in it. We'll add a negative lookahead at the start to assert the string doesn't contain the forbidden substring, then match your original structure:
^(?!.*\sda\scui\s).*\]\s1\s([\w\s]+)\s2\s([\w\s]+)$
(?!.*\sda\scui\s): This negative lookahead checks that nowhere in the string do we findda cui(surrounded by spaces, so we don't accidentally match partial hits likedacui).- The rest of the pattern follows your original logic to capture the words after
1and2.
Option 2: Restrict the captured groups from containing da cui
If you only care that the words you're capturing (after 1 and 2) don't include da cui (but the rest of the string could have it, though that's unlikely for your use case), you can modify the capture groups to avoid the forbidden substring:
\]\s1\s((?!\sda\scui\s)[\w\s])+\s2\s((?!\sda\scui\s)[\w\s])+
((?!\sda\scui\s)[\w\s])+: This replaces your original[\w\s]+. For every character we match (word or space), we first check that we're not about to hitda cui— so the captured groups never include that substring.
Example usage in Python
Here's how you'd implement Option 1 in code:
import re # Use the first pattern for full-line exclusion pattern = r'^(?!.*\sda\scui\s).*\]\s1\s([\w\s]+)\s2\s([\w\s]+)$' test_strings = [ "] 1 apple banana 2 cherry date", # Should match "] 1 hello da cui world 2 foo bar", # Should be excluded "] 1 test 2 da cui example", # Should be excluded "] 1 sun moon 2 star cloud" # Should match ] for s in test_strings: match = re.match(pattern, s) if match: print(f"Matched! Group 1: {match.group(1)}, Group 2: {match.group(2)}") else: print(f"Rejected: {s}")
Quick notes
- If you don't need to strictly match
da cuiwith spaces around it (e.g., you want to excludedacuiorda_cuitoo), remove the\saroundda\scuiin the lookahead. - If you're matching across multiple lines, add the
re.MULTILINEflag so^and$target each line's start/end instead of the whole string.
内容的提问来源于stack exchange,提问作者Anna

