如何用正则表达式检测字符串是否包含列表中至少一个指定单词?
Let's break this down step by step based on your examples and the issue you're hitting with strings like "This is my t...".
First, Clarify the Matching Rules
From your test cases, we can define the exact rules we need to follow (assuming your target word list is ["lion", "test"]):
- Exact full-word match: If the string contains the target word as a standalone word (like "lion" or "test"), return
true. - Substring match (with exceptions): For words like "lion", we allow it to appear as a substring in another word but not as a prefix or suffix of that word:
- Matches: "lotion" (lion is a middle substring), "location" (same logic)
- Doesn't match: "dandelion" (lion is a suffix), "lioness" (lion is a prefix)
- Strict full-word only for some words: For "test", we only allow exact full-word matches—so "testing" (where test is a prefix) returns
false.
Looking at your "testing" example, that makes sense: you don't want partial prefix matches for "test", but you do want non-prefix/non-suffix substring matches for "lion".
Fixing Your Code
If your existing code was missing these prefix/suffix checks, that's why you're hitting issues with "This is my t..." strings. Here are two robust solutions:
Solution 1: Manual Word Check (Easy to Customize)
Split the input string into words, then check each word against the rules:
def contains_target_word(input_str, target_list): words = input_str.split() for word in words: for target in target_list: # Rule 1: Exact full word match if word == target: return True # Rule 2: Substring match (not prefix, not suffix) # Adjust this if you only need to exclude suffixes (like for "dandelion") if target in word and not word.startswith(target) and not word.endswith(target): return True return False # Test your examples targets = ["lion", "test"] print(contains_target_word("This is my lion", targets)) # True print(contains_target_word("This is my lotion", targets)) # True print(contains_target_word("This is my dandelion", targets)) # False print(contains_target_word("This is my location", targets)) # True print(contains_target_word("This is my test", targets)) # True print(contains_target_word("This is my testing", targets)) # False
Solution 2: Regular Expressions (Cleaner for Scalable Rules)
Build regex patterns that enforce our rules, then check for matches:
import re def contains_target_word(input_str, target_list): pattern_parts = [] for target in target_list: # Exact full word match pattern_parts.append(rf'\b{re.escape(target)}\b') # Substring match (not at start or end of the word) pattern_parts.append(rf'\w+{re.escape(target)}\w+') # Combine all patterns into one regex combined_pattern = re.compile('|'.join(pattern_parts)) return bool(combined_pattern.search(input_str)) # Test the same examples targets = ["lion", "test"] print(contains_target_word("This is my lion", targets)) # True print(contains_target_word("This is my lotion", targets)) # True print(contains_target_word("This is my dandelion", targets)) # False print(contains_target_word("This is my location", targets)) # True print(contains_target_word("This is my test", targets)) # True print(contains_target_word("This is my testing", targets)) # False
Why Your Original Code Failed
If you were using a simple substring check (like target in input_str), it would incorrectly return true for "testing" (since "test" is a substring) and "dandelion" (since "lion" is a substring). Adding the prefix/suffix exclusion fixes these edge cases.
Feel free to tweak the rules if I got something slightly off—for example, if you only need to exclude suffixes (not prefixes) for "lion", just remove the not word.startswith(target) check in Solution 1, or adjust the regex pattern to rf'\w*{re.escape(target)}\w+'.
内容的提问来源于stack exchange,提问作者MalcolmInTheCenter

