匹配子字符串的替代方案:多匹配键下str.contains()的优化方法
Great question! Chaining multiple str.contains() with | gets messy real fast when you have a lot of matching keys—totally feel your pain. Let's walk through a few cleaner, more maintainable alternatives for your use case, using your sample DataFrame and mapping dictionary.
First, let's recap your setup for clarity:
import pandas as pd from pandas import DataFrame df = DataFrame({'A':['Cat had a nap','Dog had puppies','Did you see a Donkey','kitten got angry','puppy was cute']}) dic = {'Cat':'Cat','kitten':'Cat','Dog':'Dog','puppy':'Dog'}
1. Use str.extract() + map() (Vectorized & Clean)
This is my go-to for this scenario—it's efficient, concise, and scales well even with lots of keys. Here's how it works:
- Build a regex pattern by joining all your dictionary keys with
|(which acts as an OR in regex). - Extract the first matching key from each string.
- Map the extracted key to its target value using your dictionary.
# Build regex pattern from dictionary keys match_pattern = '|'.join(dic.keys()) # Extract matching key, then map to target category df['Match'] = df['A'].str.extract(f'({match_pattern})', expand=False).map(dic)
Result snippet:
| A | Match |
|---|---|
| Cat had a nap | Cat |
| Dog had puppies | Dog |
| kitten got angry | Cat |
Why this works: Vectorized pandas methods like str.extract() are way faster than row-wise operations, and the code stays clean no matter how many keys you add to your dictionary.
2. Custom apply() Function (Readable & Flexible)
If you need more control over matching logic (like prioritizing certain keys, adding exclusion rules, or handling edge cases), a custom function with apply() is perfect. It's super readable for anyone else looking at your code.
def get_matching_category(text): # Loop through key-value pairs in the dictionary for key, category in dic.items(): if key in text: return category # Return None/NaN if no match is found return pd.NA df['Match'] = df['A'].apply(get_matching_category)
Why this works: The logic is explicit—you can easily tweak the function later (e.g., change the order of checks to prioritize longer keys first) without rewriting a messy regex. The tradeoff is that apply() is slower than vectorized methods for large datasets, but it's negligible for most use cases.
3. Group Keys & Generate Boolean Columns (For Your Original Use Case)
If your goal is to create separate boolean columns (like your original df['Cat'] example), grouping keys by their target category makes this way cleaner than chaining str.contains() calls.
# Group keys by their target category first category_groups = {} for key, category in dic.items(): category_groups.setdefault(category, []).append(key) # Create a boolean column for each category for category, keys in category_groups.items(): df[category] = df['A'].str.contains('|'.join(keys))
Result snippet:
| A | Cat | Dog |
|---|---|---|
| Cat had a nap | True | False |
| puppy was cute | False | True |
Why this works: No more manual | chains—add a new key to your dictionary, and the code automatically updates the boolean columns. It's maintainable and scalable.
4. Use str.findall() + apply() (Alternative for Multiple Matches)
If you might have multiple matching keys in a single string and want to handle that (e.g., return all matching categories), str.findall() is useful:
df['All_Matches'] = df['A'].str.findall(match_pattern).apply(lambda x: [dic[key] for key in x] if x else pd.NA)
For example, if you had a string like "Cat and puppy played", this would return ['Cat', 'Dog'].
内容的提问来源于stack exchange,提问作者Sharvari Gc

