You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Window 10+Python 3.6环境下迭代查找Pandas DataFrame重复记录

Solution for Finding & Counting Duplicate (name, zip) Records

Hey there! Let's work through your problem step by step. First, let's recap your setup: you're on Windows 10 with Python 3.6, and you have this Pandas DataFrame:

import pandas as pd
df = pd.DataFrame({'name':['boo', 'foo', 'too', 'boo', 'roo', 'too'], 'zip':['30004', '02895', '02895', '30750', '02895', '02895']})

You need to find rows where both name and zip are duplicates, then count how many times each pair repeats (excluding the original entry, which matches your expected output).


Option 1: Pandas Built-in Groupby (Most Efficient for Large Data)

For large datasets, this is the best approach—Pandas' groupby is optimized under the hood, way faster than manual iteration. Here's how to do it:

# Group by both 'name' and 'zip', calculate total occurrences per group
grouped_counts = df.groupby(['name', 'zip']).size().reset_index(name='total_occurrences')

# Calculate repeat count: total occurrences minus 1 (since the first entry isn't a repeat)
grouped_counts['repeat'] = grouped_counts['total_occurrences'] - 1

# Filter to only keep groups with actual repeats, then reorder columns to match your expected output
final_result = grouped_counts[grouped_counts['repeat'] >= 1][['name', 'repeat', 'zip']].reset_index(drop=True)

print(final_result)

Output:

name  repeat    zip
0   too       1  02895

Option 2: Iterative Approach (As Requested)

If you specifically need an iterative method (e.g., for custom row-level logic), you can use a dictionary to track counts as you loop through the DataFrame:

from collections import defaultdict

# Initialize a dictionary to track how many times each (name, zip) pair appears
pair_counter = defaultdict(int)

# First pass: count all occurrences of each (name, zip) pair
for _, row in df.iterrows():
    pair_key = (row['name'], row['zip'])
    pair_counter[pair_key] += 1

# Second pass: build the result list with only repeated pairs
repeat_results = []
for (name, zip_code), count in pair_counter.items():
    if count > 1:
        repeat_results.append({
            'name': name,
            'repeat': count - 1,
            'zip': zip_code
        })

# Convert the result list to a DataFrame
final_result_df = pd.DataFrame(repeat_results).reset_index(drop=True)

print(final_result_df)

This will give you the same output as the groupby method. Note that iterrows() can be slow for very large datasets, so stick with the groupby method unless you have a specific need for iteration.


内容的提问来源于stack exchange,提问作者datanew

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:24:35