在Window 10+Python 3.6环境下迭代查找Pandas DataFrame重复记录
Hey there! Let's work through your problem step by step. First, let's recap your setup: you're on Windows 10 with Python 3.6, and you have this Pandas DataFrame:
import pandas as pd df = pd.DataFrame({'name':['boo', 'foo', 'too', 'boo', 'roo', 'too'], 'zip':['30004', '02895', '02895', '30750', '02895', '02895']})
You need to find rows where both name and zip are duplicates, then count how many times each pair repeats (excluding the original entry, which matches your expected output).
Option 1: Pandas Built-in Groupby (Most Efficient for Large Data)
For large datasets, this is the best approach—Pandas' groupby is optimized under the hood, way faster than manual iteration. Here's how to do it:
# Group by both 'name' and 'zip', calculate total occurrences per group grouped_counts = df.groupby(['name', 'zip']).size().reset_index(name='total_occurrences') # Calculate repeat count: total occurrences minus 1 (since the first entry isn't a repeat) grouped_counts['repeat'] = grouped_counts['total_occurrences'] - 1 # Filter to only keep groups with actual repeats, then reorder columns to match your expected output final_result = grouped_counts[grouped_counts['repeat'] >= 1][['name', 'repeat', 'zip']].reset_index(drop=True) print(final_result)
Output:
name repeat zip 0 too 1 02895
Option 2: Iterative Approach (As Requested)
If you specifically need an iterative method (e.g., for custom row-level logic), you can use a dictionary to track counts as you loop through the DataFrame:
from collections import defaultdict # Initialize a dictionary to track how many times each (name, zip) pair appears pair_counter = defaultdict(int) # First pass: count all occurrences of each (name, zip) pair for _, row in df.iterrows(): pair_key = (row['name'], row['zip']) pair_counter[pair_key] += 1 # Second pass: build the result list with only repeated pairs repeat_results = [] for (name, zip_code), count in pair_counter.items(): if count > 1: repeat_results.append({ 'name': name, 'repeat': count - 1, 'zip': zip_code }) # Convert the result list to a DataFrame final_result_df = pd.DataFrame(repeat_results).reset_index(drop=True) print(final_result_df)
This will give you the same output as the groupby method. Note that iterrows() can be slow for very large datasets, so stick with the groupby method unless you have a specific need for iteration.
内容的提问来源于stack exchange,提问作者datanew

