You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何清洗DataFrame单元格内的重复姓名数据?

Fixing Duplicated Names in Pandas DataFrame

Hey there! This is a super common hiccup when scraping tables that include hidden elements (like those images you noted throwing off the parsing). Let's walk through a few reliable ways to clean up those duplicated names like MarkMark, SarahSarah, or BenBen so each cell has just a single instance of the name.

Method 1: String Slicing (For Exact Double Repeats)

If every duplicated name is exactly twice the original (no extra characters, consistent length), the simplest fix is to slice the string to take the first half. This works perfectly for your sample data.

import pandas as pd

# Sample DataFrame
df = pd.DataFrame({'name': ['MarkMark', 'SarahSarah', 'BenBen']})

# Clean the names by taking the first half of each string
df['cleaned_name'] = df['name'].apply(lambda x: x[:len(x)//2])

print(df)

Output:

name cleaned_name
0    MarkMark         Mark
1  SarahSarah         Sarah
2      BenBen           Ben

Method 2: Regular Expressions (More Flexible)

If you need to handle edge cases (like accidental extra spaces or slight variations, though your example doesn't have these), regex is a great tool. This pattern matches any word that's repeated back-to-back and replaces it with a single instance.

import re

def clean_duplicate(name):
    # Regex to match a word that's repeated exactly at the start/end
    return re.sub(r'^(\w+)\1$', r'\1', name)

df['cleaned_name'] = df['name'].apply(clean_duplicate)

The regex ^(\w+)\1$ breaks down as:

  • ^ = start of string
  • (\w+) = capture group for one or more word characters (letters, numbers, underscores)
  • \1 = matches the exact same text as the first capture group
  • $ = end of string

Method 3: Conditional Cleaning (Avoid Breaking Valid Names)

If your dataset has a mix of duplicated names and normal ones, add a check to only modify strings that are true duplicates (to avoid accidentally truncating valid names):

def safe_clean_name(name):
    name_length = len(name)
    # Only proceed if the string length is even, and first half matches second half
    if name_length % 2 == 0:
        half_point = name_length // 2
        if name[:half_point] == name[half_point:]:
            return name[:half_point]
    # Return original name if it doesn't meet the duplicate criteria
    return name

df['cleaned_name'] = df['name'].apply(safe_clean_name)

This way, if you have a valid name like Anna (which isn't a duplicate of anything), it stays unchanged instead of getting sliced to An.

Quick Validation Step

After cleaning, double-check that the fix worked for all cases:

# Verify that cleaned name doubled equals original (for duplicates)
duplicate_rows = df[df['name'] != df['cleaned_name']]
print(all(duplicate_rows['cleaned_name'].apply(lambda x: x*2) == duplicate_rows['name']))

If this returns True, your cleaning worked as expected!

内容的提问来源于stack exchange,提问作者Matt S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 10:43:12