You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Fuzzy Wuzzy替代Regex实现Pandas列的模糊匹配并处理拼写错误?

Replacing Regex with Fuzzy Matching (and Other NLP Solutions) for Scalable Text Classification

First, let's fix the core problem: replicating your regex logic (^(?=.*SEA)(?=.*MAN)(?!.*WOMAN)(?!.*CHILD)) with fuzzy matching to handle typos, overlapping terms, and scalable keyword management. We'll use RapidFuzz (a maintained, faster alternative to Fuzzy Wuzzy) for string similarity checks, then expand to more robust NLP approaches.


Step 1: Basic Fuzzy Matching to Replicate Regex Logic

Your regex checks for two required terms (SEA, MAN) and excludes two forbidden terms (WOMAN, CHILD). With fuzzy matching, we'll adjust this to account for spelling errors and partial matches (like "seaman" containing both SEA and MAN).

Setup

First install dependencies:

pip install rapidfuzz pandas

Helper Functions for Inclusion/Exclusion Checks

These functions use partial_ratio to detect if a term is present (even as part of a longer word) with a configurable similarity threshold:

import pandas as pd
from rapidfuzz import fuzz

def meets_required_terms(text, required_terms, threshold=80):
    """Check if all required terms are fuzzy-matched in the text"""
    text_lower = text.lower()
    for term in required_terms:
        term_lower = term.lower()
        # partial_ratio checks if the term is a fuzzy substring of the text
        if fuzz.partial_ratio(text_lower, term_lower) < threshold:
            return False
    return True

def avoids_forbidden_terms(text, forbidden_terms, threshold=80):
    """Check if no forbidden terms are fuzzy-matched in the text"""
    text_lower = text.lower()
    for term in forbidden_terms:
        term_lower = term.lower()
        if fuzz.partial_ratio(text_lower, term_lower) >= threshold:
            return False
    return True

Apply to Your DataFrame

Replicate your original regex logic with fuzzy checks:

# Example dataframe
df = pd.DataFrame({
    'random_words': ['sea man', 'seaman', 'woman at sea', 'child and boat', 'sailor man']
})

# Assign 'Fisherman' where rules are met
df.loc[
    df['random_words'].apply(lambda x: meets_required_terms(x, ['SEA', 'MAN']) and avoids_forbidden_terms(x, ['WOMAN', 'CHILD'])),
    'Person'
] = 'Fisherman'

# Output: rows 0,1,4 get 'Fisherman'
print(df)

Step 2: Scale to 50+ Category Rules

To manage 50 fixed result keywords, create a dictionary of rules for each category, then use a scoring system to pick the best match:

Define Category Rules

category_rules = {
    'Fisherman': {
        'required': ['SEA', 'MAN', 'SAIL'],
        'forbidden': ['WOMAN', 'CHILD', 'DOCTOR']
    },
    'Nurse': {
        'required': ['HOSPITAL', 'CARE', 'PATIENT'],
        'forbidden': ['MECHANIC', 'CAR', 'BOAT']
    },
    # Add all 50 categories here...
}

Score and Assign Categories

This function calculates a score for each category and returns the highest-scoring valid match:

def calculate_category_score(text, rule, threshold=80):
    text_lower = text.lower()
    total_score = 0
    
    # Add points for required terms
    for term in rule['required']:
        ratio = fuzz.partial_ratio(text_lower, term.lower())
        if ratio < threshold:
            return -float('inf')  # Fail required terms
        total_score += ratio
    
    # Subtract points for forbidden terms (or invalidate if threshold is hit)
    for term in rule['forbidden']:
        ratio = fuzz.partial_ratio(text_lower, term.lower())
        if ratio >= threshold:
            return -float('inf')  # Forbidden term found
        total_score -= ratio
    
    return total_score

def assign_best_category(text, rules, threshold=80):
    scores = {}
    for category, rule in rules.items():
        score = calculate_category_score(text, rule, threshold)
        if score > -float('inf'):
            scores[category] = score
    
    if not scores:
        return 'Unknown'
    # Return category with highest score
    return max(scores, key=scores.get)

# Apply to dataframe
df['Person'] = df['random_words'].apply(assign_best_category, rules=category_rules)

Step 3: Fixing Your Existing Matching Function

If you were trying to match rows between two dataframes (df1[A] and df2[B]) and assign the best match, use RapidFuzz's extractOne function:

from rapidfuzz import process

def get_best_cross_match(text, choices, threshold=80):
    match, score = process.extractOne(text, choices)
    return match if score >= threshold else None

# Example: Match df1['A'] to best term in df2['B']
df1['best_match'] = df1['A'].apply(get_best_cross_match, choices=df2['B'].tolist())

Advanced NLP: Semantic Matching with Word Embeddings

For cases where you need to capture synonyms (e.g., "ocean" = "sea") rather than just string similarity, use pre-trained word embeddings with spaCy:

Setup

pip install spacy
python -m spacy download en_core_web_md

Semantic Matching Code

import spacy
from sklearn.metrics.pairwise import cosine_similarity

nlp = spacy.load('en_core_web_md')

def get_embedding(text):
    return nlp(text).vector

# Precompute embeddings for your rules
category_rules_emb = {
    'Fisherman': {
        'required': [get_embedding('sea'), get_embedding('man')],
        'forbidden': [get_embedding('woman'), get_embedding('child')]
    }
}

def meets_semantic_requirements(text, required_embs, threshold=0.7):
    text_emb = get_embedding(text)
    for emb in required_embs:
        sim = cosine_similarity([text_emb], [emb])[0][0]
        if sim < threshold:
            return False
    return True

def avoids_semantic_forbidden(text, forbidden_embs, threshold=0.7):
    text_emb = get_embedding(text)
    for emb in forbidden_embs:
        sim = cosine_similarity([text_emb], [emb])[0][0]
        if sim >= threshold:
            return False
    return True

# Apply to dataframe
df.loc[
    df['random_words'].apply(lambda x: meets_semantic_requirements(x, category_rules_emb['Fisherman']['required']) and avoids_semantic_forbidden(x, category_rules_emb['Fisherman']['forbidden'])),
    'Person'
] = 'Fisherman'

This approach captures contextual similarity, making it ideal for handling nuanced document descriptions.


内容的提问来源于stack exchange,提问作者Wolfy14

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 12:32:43