Python文本清洗求助:提取首项国籍与修复姓名特殊字符
Hey there! Let's work through these two data cleaning challenges together—they're super typical when dealing with real-world datasets, so I’ve got you covered.
For entries with dual nationalities separated by / (like France/America), the easiest way to grab the first country is using pandas' string splitting capabilities. Here's how you can do it:
First, assuming your dataset is loaded into a pandas DataFrame called df:
import pandas as pd # Split the nationality string by '/' and take the first element df['cleaned_nationality'] = df['nationality'].str.split('/').str[0] # Handle any missing values (optional but recommended) df['cleaned_nationality'] = df['cleaned_nationality'].fillna('Unknown')
This works because str.split('/') turns each entry into a list (e.g., ['France', 'America']), and str[0] pulls out the first item in that list. For entries without a /, it just returns the original string, which is exactly what we want.
The weird characters like ̦ in Gy̦rgy Kepes are almost always encoding mismatches (usually from mixing UTF-8 and Latin-1 encodings). There are two solid approaches to fix this:
Option 1: Use Python's built-in unicodedata module
This is great if you want to stick to standard libraries. The NFKC normalization will strip out those odd combining characters:
import unicodedata def clean_special_chars(name): # Normalize to remove combining diacritics and odd symbols normalized = unicodedata.normalize('NFKC', name) # Optional: Remove any remaining non-printable characters return ''.join(c for c in normalized if c.isprintable()) # Apply the function to your name column df['cleaned_artist_name'] = df['artist_name'].apply(clean_special_chars)
Option 2: Use the ftfy library (more powerful for tricky cases)
If the built-in method doesn't fix everything, ftfy is a dedicated library for fixing broken text encoding. First install it via pip:
pip install ftfy
Then use it like this:
from ftfy import fix_text df['cleaned_artist_name'] = df['artist_name'].apply(fix_text)
ftfy is especially good at detecting and reversing common encoding mistakes (like when text was decoded with the wrong charset), so it’ll likely turn Gy̦rgy Kepes into the correct György Kepes.
Pro Tip: Check the Source Encoding First
Before cleaning, make sure you loaded the dataset with the correct encoding. If you used pd.read_csv, try specifying encoding='latin-1' or encoding='utf-8' to see if that prevents the special characters from appearing in the first place:
df = pd.read_csv('your_dataset.csv', encoding='latin-1')
内容的提问来源于stack exchange,提问作者Alonso Ag

