You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否在DataFrame列中使用Fuzzy Wuzzy?求DataFrame计数列实现方案

Hey there! Let's break down your questions and solve them step by step.

1. Adding the count Column to Your DataFrame

Your goal is to increment a counter whenever rows share the same col1 and col2 values, following the original row order. This is perfect for pandas' groupby() combined with cumcount()—here's how to do it:

First, let's recreate your original DataFrame to test with:

import pandas as pd

data = {
    'col1': ['A1', 'A1', 'A2', 'A3', 'A2', 'A4', 'A1'],
    'col2': ['B1', 'B1', 'B1', 'B2', 'B1', 'B3', 'B1'],
    'col3': ['C1', 'C2', 'C3', 'C4', 'C5', 'C6', 'C1']
}
df = pd.DataFrame(data)

Now, add the count column by grouping on col1 and col2, then using cumcount() (which starts at 0, so we add 1 to get your desired 1-based counter):

df['count'] = df.groupby(['col1', 'col2']).cumcount() + 1

If you run this, your resulting DataFrame will match exactly what you wanted:

col1 col2 col3  count
0   A1   B1   C1      1
1   A1   B1   C2      2
2   A2   B1   C3      1
3   A3   B2   C4      1
4   A2   B1   C5      2
5   A4   B3   C6      1
6   A1   B1   C1      3

The cumcount() function tracks the position of each row within its (col1, col2) group, so it increments exactly as you need it to.

2. Using Fuzzy Wuzzy with DataFrames

Absolutely! You can use Fuzzy Wuzzy to perform string matching operations directly on DataFrame columns. First, make sure you have it installed (along with its dependency python-Levenshtein for better performance):

pip install fuzzywuzzy python-Levenshtein

Here are a few common use cases to get you started:

  • Calculate Similarity Between Columns
    If you want to compare strings in col3 with a fixed string or another column, use apply() with Fuzzy Wuzzy's fuzz.ratio():

    from fuzzywuzzy import fuzz
    
    # Compare each value in col3 to the string "C1"
    df['similarity_to_C1'] = df['col3'].apply(lambda x: fuzz.ratio(x, 'C1'))
    
  • Find Most Similar String in a Group
    To find the closest matching col3 value within each (col1, col2) group, combine groupby() with process.extractOne():

    from fuzzywuzzy import process
    
    def find_most_similar(group):
        values = group['col3'].tolist()
        return group['col3'].apply(lambda x: process.extractOne(x, [v for v in values if v != x])[0])
    
    df['most_similar_col3'] = df.groupby(['col1', 'col2']).apply(find_most_similar).reset_index(drop=True)
    
  • Filter Rows by Similarity Threshold
    You can filter rows where two columns have a similarity score above a certain threshold:

    # Add a test column for demonstration
    df['col3_test'] = ['C1', 'C2', 'D3', 'C4', 'D5', 'D6', 'C1']
    
    # Keep rows where col3 and col3_test have a similarity score ≥ 80
    filtered_df = df[df.apply(lambda row: fuzz.ratio(row['col3'], row['col3_test']) >= 80, axis=1)]
    

Just note that Fuzzy Wuzzy operations can be slower on large DataFrames since they're string-based. For very big datasets, you might want to use rapidfuzz—a faster drop-in alternative to Fuzzy Wuzzy.

内容的提问来源于stack exchange,提问作者user7490407

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:25:03