如何判断Pandas Series中的字符串是否包含另一个DataFrame中的字符串?
Hey there! Let's work through this problem together. You've got two DataFrames: one with animal descriptions (desc_df) and another with lists of cat and dog breeds (catg_df), and you want to check which descriptions include any of those breeds. Here's how to do it step by step:
Step 1: Extract Clean Breed Lists from catg_df
First, let's pull out the actual breed names from catg_df—we'll drop any empty values (in case there are NaNs) and turn each column into a simple, usable list:
import pandas as pd # Your original data setup data = {'col1': ['black sphynx bob','brown labrador','grey labrador mervin', 'brown siamese cat','white siamese']} desc_df = pd.DataFrame(data=data) catg = {'dog': ['labrador','rottweiler', 'beagle'],'cat':['siamese','sphynx','ragdoll']} catg_df = pd.DataFrame(data=catg) # Extract breed lists, removing any missing values dog_breeds = catg_df['dog'].dropna().tolist() cat_breeds = catg_df['cat'].dropna().tolist()
Step 2: Flag Breed Matches in desc_df
We can use Pandas' str.contains() method with a regex pattern to check if any breed appears in each description. Regex lets us match multiple terms at once using the | (OR) operator. We'll add case=False to ignore uppercase/lowercase differences (so "Siamese" and "siamese" both get matched):
# Build regex patterns for dog and cat breeds dog_pattern = '|'.join(dog_breeds) cat_pattern = '|'.join(cat_breeds) # Add new columns to flag if a dog/cat breed is present desc_df['contains_dog'] = desc_df['col1'].str.contains(dog_pattern, case=False) desc_df['contains_cat'] = desc_df['col1'].str.contains(cat_pattern, case=False)
What the Result Looks Like
After running that code, your desc_df will have clear flags for each row:
| col1 | contains_dog | contains_cat |
|---|---|---|
| black sphynx bob | False | True |
| brown labrador | True | False |
| grey labrador mervin | True | False |
| brown siamese cat | False | True |
| white siamese | False | True |
Bonus: Get the Exact Matching Breed
If you want to know which breed was matched (not just if there's a match), use a custom function with apply() to pull out the specific breed name:
def find_matching_breeds(text, breed_list): # Check each breed, ignoring case differences matches = [breed for breed in breed_list if breed.lower() in text.lower()] return ', '.join(matches) if matches else None # Add columns showing exactly which breeds were found desc_df['matching_dog_breeds'] = desc_df['col1'].apply(lambda x: find_matching_breeds(x, dog_breeds)) desc_df['matching_cat_breeds'] = desc_df['col1'].apply(lambda x: find_matching_breeds(x, cat_breeds))
This will add columns like matching_cat_breeds that show "sphynx" for the first row, or "siamese" for the fourth row.
Quick Spelling Note
I noticed a small typo in your sample desc_df ("black spyhnx bob" instead of "sphynx")—that would fail to match unless you fix the spelling. If you need to handle typos, you'd need fuzzy matching (like using the fuzzywuzzy library), but for exact matches, double-checking spelling is key.
内容的提问来源于stack exchange,提问作者CGully

