如何在DataFrame行中返回字典匹配值对应的多个键(以|分隔)
Got it, let's fix this up so you can capture all matching keys per row and join them with a |! Here's a straightforward approach using pandas and a custom function:
Step 1: Define Your Target Vocabulary Dictionary
First, lay out the key-value pairs you're looking to extract. Let's use a category-based example that aligns with your sample product names:
# Replace this with your actual target vocabulary and mappings target_map = { 'Red': 'warm_color', 'Blue': 'cool_color', 'Azure': 'cool_color', 'Ruby': 'warm_color', 'Lace': 'fabric_detail', 'Sweater': 'top_garment' }
Step 2: Build a Custom Matching Function
This function will scan each string, collect all matching values from your dictionary, and join them with |. If no matches are found, it returns np.nan (you can swap this for an empty string if preferred):
def extract_all_matches(text): # Gather all values where the corresponding key exists in the input text matched_values = [value for key, value in target_map.items() if key in text] # Join matches with | if any exist, else return NaN return '|'.join(matched_values) if matched_values else np.nan
Step 3: Apply the Function to Your DataFrame
Use pandas' apply() method to run this function across every row in your target column:
import pandas as pd import numpy as np # Your sample DataFrame df = pd.DataFrame({ 'Name': [ 'Red and Blue Lace Midi Dress', 'Long Armed Sweater Azure and Ruby', 'High Waisted Cotton Pants' # Added a test row with no matches ] }) # Add a new column with the combined matches df['Matched_Categories'] = df['Name'].apply(extract_all_matches) # Print the result print(df)
Sample Output
Name Matched_Categories 0 Red and Blue Lace Midi Dress warm_color|cool_color|fabric_detail 1 Long Armed Sweater Azure and Ruby warm_color|cool_color|top_garment 2 High Waisted Cotton Pants NaN
Bonus: Whole-Word Matching (Avoid Partial Hits)
If you want to prevent partial matches (e.g., not picking up "Red" in "Reddish"), use regex with word boundaries:
import re def extract_all_matches(text): matched_values = [] for key, value in target_map.items(): # Use regex to match whole words only if re.search(rf'\b{re.escape(key)}\b', text): matched_values.append(value) return '|'.join(matched_values) if matched_values else np.nan
内容的提问来源于stack exchange,提问作者Bob Harris

