如何清洗DataFrame单元格内的重复姓名数据?
Hey there! This is a super common hiccup when scraping tables that include hidden elements (like those images you noted throwing off the parsing). Let's walk through a few reliable ways to clean up those duplicated names like MarkMark, SarahSarah, or BenBen so each cell has just a single instance of the name.
Method 1: String Slicing (For Exact Double Repeats)
If every duplicated name is exactly twice the original (no extra characters, consistent length), the simplest fix is to slice the string to take the first half. This works perfectly for your sample data.
import pandas as pd # Sample DataFrame df = pd.DataFrame({'name': ['MarkMark', 'SarahSarah', 'BenBen']}) # Clean the names by taking the first half of each string df['cleaned_name'] = df['name'].apply(lambda x: x[:len(x)//2]) print(df)
Output:
name cleaned_name 0 MarkMark Mark 1 SarahSarah Sarah 2 BenBen Ben
Method 2: Regular Expressions (More Flexible)
If you need to handle edge cases (like accidental extra spaces or slight variations, though your example doesn't have these), regex is a great tool. This pattern matches any word that's repeated back-to-back and replaces it with a single instance.
import re def clean_duplicate(name): # Regex to match a word that's repeated exactly at the start/end return re.sub(r'^(\w+)\1$', r'\1', name) df['cleaned_name'] = df['name'].apply(clean_duplicate)
The regex ^(\w+)\1$ breaks down as:
^= start of string(\w+)= capture group for one or more word characters (letters, numbers, underscores)\1= matches the exact same text as the first capture group$= end of string
Method 3: Conditional Cleaning (Avoid Breaking Valid Names)
If your dataset has a mix of duplicated names and normal ones, add a check to only modify strings that are true duplicates (to avoid accidentally truncating valid names):
def safe_clean_name(name): name_length = len(name) # Only proceed if the string length is even, and first half matches second half if name_length % 2 == 0: half_point = name_length // 2 if name[:half_point] == name[half_point:]: return name[:half_point] # Return original name if it doesn't meet the duplicate criteria return name df['cleaned_name'] = df['name'].apply(safe_clean_name)
This way, if you have a valid name like Anna (which isn't a duplicate of anything), it stays unchanged instead of getting sliced to An.
Quick Validation Step
After cleaning, double-check that the fix worked for all cases:
# Verify that cleaned name doubled equals original (for duplicates) duplicate_rows = df[df['name'] != df['cleaned_name']] print(all(duplicate_rows['cleaned_name'].apply(lambda x: x*2) == duplicate_rows['name']))
If this returns True, your cleaning worked as expected!
内容的提问来源于stack exchange,提问作者Matt S

