如何在Python中检测字符串是否包含字典内的指定词汇?
Got it, let's break this down. You have a dictionary where one key maps to a list of keywords, and you want to check if an input string (like "I love simplicity") includes any of those keywords. Here are two reliable ways to implement this:
Method 1: Clean and Split the Input String
This approach works well for basic cases, where you want to match exact words after cleaning up punctuation and case differences:
import string # Your dictionary data review_keywords = {"value for money": ["rescheduled", "cost", "low", "high", "simplicity", "booking", "price-performance", "satisfied", "satisfaction", "pricing", "prices"]} # Extract the list of target keywords target_words = review_keywords["value for money"] def contains_target_keyword(input_str): # Step 1: Clean the input string - lowercase and remove punctuation cleaned_str = input_str.lower().translate(str.maketrans('', '', string.punctuation)) # Step 2: Split into individual words input_words = cleaned_str.split() # Step 3: Check if any target word exists in the input words return any(word in input_words for word in target_words) # Test cases print(contains_target_keyword("I love simplicity")) # Output: True print(contains_target_keyword("The price is way too high!")) # Output: True print(contains_target_keyword("This service is amazing")) # Output: False
Key Notes for Method 1:
- We convert the input to lowercase to avoid missing matches like "Simplicity" vs "simplicity".
- We strip punctuation so words like "simplicity!" are treated the same as "simplicity".
any()stops checking as soon as it finds a match, making it efficient.
Method 2: Regular Expression (More Robust)
If you want to ensure whole-word matches (so "lower" doesn't trigger a match for "low") and handle edge cases like words with special characters (e.g., "price-performance"), regex is a better choice:
import re # Your dictionary data review_keywords = {"value for money": ["rescheduled", "cost", "low", "high", "simplicity", "booking", "price-performance", "satisfied", "satisfaction", "pricing", "prices"]} target_words = review_keywords["value for money"] # Build a regex pattern: match any target word as a whole word, case-insensitive # re.escape() handles special characters in keywords (like the hyphen in price-performance) pattern = re.compile( r'\b(' + '|'.join(re.escape(word) for word in target_words) + r')\b', re.IGNORECASE ) def contains_target_keyword(input_str): # Check if the pattern exists anywhere in the input string return bool(pattern.search(input_str)) # Test cases print(contains_target_keyword("I love Simplicity!")) # Output: True print(contains_target_keyword("The lower price is great")) # Output: False (doesn't match "low") print(contains_target_keyword("Great price-performance on this booking")) # Output: True
Key Notes for Method 2:
\bensures we only match whole words (not substrings).re.IGNORECASEmakes the match case-insensitive.re.escape()prevents special characters in keywords from breaking the regex pattern.
Either method will work for your use case—pick the one that fits your specific needs!
内容的提问来源于stack exchange,提问作者user3310469

