能否在DataFrame列中使用Fuzzy Wuzzy?求DataFrame计数列实现方案
Hey there! Let's break down your questions and solve them step by step.
count Column to Your DataFrame Your goal is to increment a counter whenever rows share the same col1 and col2 values, following the original row order. This is perfect for pandas' groupby() combined with cumcount()—here's how to do it:
First, let's recreate your original DataFrame to test with:
import pandas as pd data = { 'col1': ['A1', 'A1', 'A2', 'A3', 'A2', 'A4', 'A1'], 'col2': ['B1', 'B1', 'B1', 'B2', 'B1', 'B3', 'B1'], 'col3': ['C1', 'C2', 'C3', 'C4', 'C5', 'C6', 'C1'] } df = pd.DataFrame(data)
Now, add the count column by grouping on col1 and col2, then using cumcount() (which starts at 0, so we add 1 to get your desired 1-based counter):
df['count'] = df.groupby(['col1', 'col2']).cumcount() + 1
If you run this, your resulting DataFrame will match exactly what you wanted:
col1 col2 col3 count 0 A1 B1 C1 1 1 A1 B1 C2 2 2 A2 B1 C3 1 3 A3 B2 C4 1 4 A2 B1 C5 2 5 A4 B3 C6 1 6 A1 B1 C1 3
The cumcount() function tracks the position of each row within its (col1, col2) group, so it increments exactly as you need it to.
Absolutely! You can use Fuzzy Wuzzy to perform string matching operations directly on DataFrame columns. First, make sure you have it installed (along with its dependency python-Levenshtein for better performance):
pip install fuzzywuzzy python-Levenshtein
Here are a few common use cases to get you started:
Calculate Similarity Between Columns
If you want to compare strings incol3with a fixed string or another column, useapply()with Fuzzy Wuzzy'sfuzz.ratio():from fuzzywuzzy import fuzz # Compare each value in col3 to the string "C1" df['similarity_to_C1'] = df['col3'].apply(lambda x: fuzz.ratio(x, 'C1'))Find Most Similar String in a Group
To find the closest matchingcol3value within each(col1, col2)group, combinegroupby()withprocess.extractOne():from fuzzywuzzy import process def find_most_similar(group): values = group['col3'].tolist() return group['col3'].apply(lambda x: process.extractOne(x, [v for v in values if v != x])[0]) df['most_similar_col3'] = df.groupby(['col1', 'col2']).apply(find_most_similar).reset_index(drop=True)Filter Rows by Similarity Threshold
You can filter rows where two columns have a similarity score above a certain threshold:# Add a test column for demonstration df['col3_test'] = ['C1', 'C2', 'D3', 'C4', 'D5', 'D6', 'C1'] # Keep rows where col3 and col3_test have a similarity score ≥ 80 filtered_df = df[df.apply(lambda row: fuzz.ratio(row['col3'], row['col3_test']) >= 80, axis=1)]
Just note that Fuzzy Wuzzy operations can be slower on large DataFrames since they're string-based. For very big datasets, you might want to use rapidfuzz—a faster drop-in alternative to Fuzzy Wuzzy.
内容的提问来源于stack exchange,提问作者user7490407

