如何使用Fuzzy Wuzzy替代Regex实现Pandas列的模糊匹配并处理拼写错误?
First, let's fix the core problem: replicating your regex logic (^(?=.*SEA)(?=.*MAN)(?!.*WOMAN)(?!.*CHILD)) with fuzzy matching to handle typos, overlapping terms, and scalable keyword management. We'll use RapidFuzz (a maintained, faster alternative to Fuzzy Wuzzy) for string similarity checks, then expand to more robust NLP approaches.
Step 1: Basic Fuzzy Matching to Replicate Regex Logic
Your regex checks for two required terms (SEA, MAN) and excludes two forbidden terms (WOMAN, CHILD). With fuzzy matching, we'll adjust this to account for spelling errors and partial matches (like "seaman" containing both SEA and MAN).
Setup
First install dependencies:
pip install rapidfuzz pandas
Helper Functions for Inclusion/Exclusion Checks
These functions use partial_ratio to detect if a term is present (even as part of a longer word) with a configurable similarity threshold:
import pandas as pd from rapidfuzz import fuzz def meets_required_terms(text, required_terms, threshold=80): """Check if all required terms are fuzzy-matched in the text""" text_lower = text.lower() for term in required_terms: term_lower = term.lower() # partial_ratio checks if the term is a fuzzy substring of the text if fuzz.partial_ratio(text_lower, term_lower) < threshold: return False return True def avoids_forbidden_terms(text, forbidden_terms, threshold=80): """Check if no forbidden terms are fuzzy-matched in the text""" text_lower = text.lower() for term in forbidden_terms: term_lower = term.lower() if fuzz.partial_ratio(text_lower, term_lower) >= threshold: return False return True
Apply to Your DataFrame
Replicate your original regex logic with fuzzy checks:
# Example dataframe df = pd.DataFrame({ 'random_words': ['sea man', 'seaman', 'woman at sea', 'child and boat', 'sailor man'] }) # Assign 'Fisherman' where rules are met df.loc[ df['random_words'].apply(lambda x: meets_required_terms(x, ['SEA', 'MAN']) and avoids_forbidden_terms(x, ['WOMAN', 'CHILD'])), 'Person' ] = 'Fisherman' # Output: rows 0,1,4 get 'Fisherman' print(df)
Step 2: Scale to 50+ Category Rules
To manage 50 fixed result keywords, create a dictionary of rules for each category, then use a scoring system to pick the best match:
Define Category Rules
category_rules = { 'Fisherman': { 'required': ['SEA', 'MAN', 'SAIL'], 'forbidden': ['WOMAN', 'CHILD', 'DOCTOR'] }, 'Nurse': { 'required': ['HOSPITAL', 'CARE', 'PATIENT'], 'forbidden': ['MECHANIC', 'CAR', 'BOAT'] }, # Add all 50 categories here... }
Score and Assign Categories
This function calculates a score for each category and returns the highest-scoring valid match:
def calculate_category_score(text, rule, threshold=80): text_lower = text.lower() total_score = 0 # Add points for required terms for term in rule['required']: ratio = fuzz.partial_ratio(text_lower, term.lower()) if ratio < threshold: return -float('inf') # Fail required terms total_score += ratio # Subtract points for forbidden terms (or invalidate if threshold is hit) for term in rule['forbidden']: ratio = fuzz.partial_ratio(text_lower, term.lower()) if ratio >= threshold: return -float('inf') # Forbidden term found total_score -= ratio return total_score def assign_best_category(text, rules, threshold=80): scores = {} for category, rule in rules.items(): score = calculate_category_score(text, rule, threshold) if score > -float('inf'): scores[category] = score if not scores: return 'Unknown' # Return category with highest score return max(scores, key=scores.get) # Apply to dataframe df['Person'] = df['random_words'].apply(assign_best_category, rules=category_rules)
Step 3: Fixing Your Existing Matching Function
If you were trying to match rows between two dataframes (df1[A] and df2[B]) and assign the best match, use RapidFuzz's extractOne function:
from rapidfuzz import process def get_best_cross_match(text, choices, threshold=80): match, score = process.extractOne(text, choices) return match if score >= threshold else None # Example: Match df1['A'] to best term in df2['B'] df1['best_match'] = df1['A'].apply(get_best_cross_match, choices=df2['B'].tolist())
Advanced NLP: Semantic Matching with Word Embeddings
For cases where you need to capture synonyms (e.g., "ocean" = "sea") rather than just string similarity, use pre-trained word embeddings with spaCy:
Setup
pip install spacy python -m spacy download en_core_web_md
Semantic Matching Code
import spacy from sklearn.metrics.pairwise import cosine_similarity nlp = spacy.load('en_core_web_md') def get_embedding(text): return nlp(text).vector # Precompute embeddings for your rules category_rules_emb = { 'Fisherman': { 'required': [get_embedding('sea'), get_embedding('man')], 'forbidden': [get_embedding('woman'), get_embedding('child')] } } def meets_semantic_requirements(text, required_embs, threshold=0.7): text_emb = get_embedding(text) for emb in required_embs: sim = cosine_similarity([text_emb], [emb])[0][0] if sim < threshold: return False return True def avoids_semantic_forbidden(text, forbidden_embs, threshold=0.7): text_emb = get_embedding(text) for emb in forbidden_embs: sim = cosine_similarity([text_emb], [emb])[0][0] if sim >= threshold: return False return True # Apply to dataframe df.loc[ df['random_words'].apply(lambda x: meets_semantic_requirements(x, category_rules_emb['Fisherman']['required']) and avoids_semantic_forbidden(x, category_rules_emb['Fisherman']['forbidden'])), 'Person' ] = 'Fisherman'
This approach captures contextual similarity, making it ideal for handling nuanced document descriptions.
内容的提问来源于stack exchange,提问作者Wolfy14

