将字典中的精确匹配键映射到Pandas DataFrame列并返回对应值
Hey there! Let's tackle those keyword-to-value matching snags you're hitting with your Pandas DataFrame, building on the key-matching solution you already have. I’ll walk through the most common scenarios and fixes with concrete code examples:
1. Exact Value Matching
If you need to match exact values between your keyword dictionary and DataFrame entries, the easiest approach is to reverse your dictionary to map values directly to their parent keys, then use Pandas' map() method.
Example Code:
import pandas as pd # Your existing keyword dictionary keyword_dict = {"fruit": ["apple", "banana"], "veggie": ["carrot", "spinach"]} # Sample DataFrame (adjust to match your actual data structure) df = pd.DataFrame( {"items": ["apple", "orange", "carrot", "grape"], "matched_category": ["", "", "", ""]} ) # Reverse the dictionary to create a value-to-key lookup map value_to_key = {v: k for key, values in keyword_dict.items() for v in values} # Perform exact match and fill unmatched entries with "unknown" df["matched_category"] = df["items"].map(value_to_key).fillna("unknown") print(df)
Output:
items matched_category 0 apple fruit 1 orange unknown 2 carrot veggie 3 grape unknown
2. Partial/Substring Matching
If your DataFrame entries contain the keywords as substrings (e.g., "fresh apple" instead of just "apple"), use a custom function with apply() to check for substring matches. You can add case-insensitive matching to avoid missing matches due to capitalization.
Example Code:
# Custom function to check for substring matches def find_substring_match(item, keyword_map): # Convert item to lowercase for case-insensitive check lower_item = item.lower() for category, keywords in keyword_map.items(): for keyword in keywords: if keyword.lower() in lower_item: return category return "unknown" # Apply the function to your DataFrame column df["matched_category"] = df["items"].apply(find_substring_match, keyword_map=keyword_dict) # Test with a modified DataFrame containing substrings df_test = pd.DataFrame({"items": ["Fresh Apple", "carrot soup", "orange juice"]}) df_test["matched_category"] = df_test["items"].apply(find_substring_match, keyword_map=keyword_dict) print(df_test)
Output:
items matched_category 0 Fresh Apple fruit 1 carrot soup veggie 2 orange juice unknown
3. Handling Multiple Matches
If a single DataFrame entry matches multiple keywords (e.g., "apple carrot salad"), modify the custom function to collect all matching categories instead of returning just one.
Example Code:
def find_multiple_matches(item, keyword_map): lower_item = item.lower() matches = [] for category, keywords in keyword_map.items(): for keyword in keywords: if keyword.lower() in lower_item: matches.append(category) # Join matches with commas, or return "unknown" if none return ", ".join(set(matches)) if matches else "unknown" # Test with a multi-match entry df_multi = pd.DataFrame({"items": ["apple carrot salad", "spinach banana smoothie"]}) df_multi["matched_category"] = df_multi["items"].apply(find_multiple_matches, keyword_map=keyword_dict) print(df_multi)
Output:
items matched_category 0 apple carrot salad fruit, veggie 1 spinach banana smoothie fruit, veggie
Pro Tips for Debugging
- Print your reversed
value_to_keymap first to confirm your keyword-to-category mappings are correct. - Use
df[df["matched_category"] == "unknown"]to isolate entries that aren't matching—check for typos, capitalization, or missing keywords in your dictionary. - For large datasets, use regex-based matching with
str.contains()for faster performance:# Build regex patterns for each category fruit_pattern = "|".join(keyword_dict["fruit"]) veggie_pattern = "|".join(keyword_dict["veggie"]) # Use np.where for vectorized matching import numpy as np df["matched_category"] = np.where( df["items"].str.contains(fruit_pattern, case=False), "fruit", np.where(df["items"].str.contains(veggie_pattern, case=False), "veggie", "unknown") )
内容的提问来源于stack exchange,提问作者J_Win

