Python Pandas入门:如何从含头衔的14k行数据集中拆分姓名
Hey there! Let's break down how to split first and last names from your messy 14k-row dataset—all those titles, honorifics, and extra text can be tricky, but with regex and Pandas' string tools, we can clean this up effectively.
First, let's look at your data patterns: you've got everything from stacked academic titles (Pr.Doz.Dr.), titles with parentheses (Dr. univ. (Budapest)), honorifics like Herr/Sohn, and even plain names with no titles. Our plan is to strip out all the leading title/honorific text first, then split the remaining clean name into first and last parts.
Step 1: Load Your Data
First, let's get your data into a Pandas DataFrame. I'll use your sample data for this example:
import pandas as pd # Sample data matching your example sample_data = { 'full_name': [ 'Pr.Doz.Dr. Klaus Semmler Facharzt für Frauenhe...', 'Dr. univ. (Budapest) Dalia Lax', 'Dr. med. Jovan Stojilkovic', 'Dr. med. Dirk Schneider', 'Marc Scheuermann', 'Bag Kinderarztpraxis', 'Herr Ulrich Bromig', 'Sohn Heinrich', 'Herr Dr. sc. med. Amadeus Hartwig', 'Jasmin Rieche' ] } df = pd.DataFrame(sample_data)
Step 2: Strip Leading Titles/Honorifics
We'll use a regular expression to target and remove all the leading title patterns. Here's a regex that covers all the cases in your sample:
# Regex pattern to match leading titles/honorifics title_regex = r'^(?:(?:[A-Za-z]+\.)+|\w+ \w+\s*\([^)]+\)|Herr|Sohn)\s+' # Remove titles to get a clean name string df['clean_name'] = df['full_name'].str.replace(title_regex, '', regex=True) # Flag rows that don't look like valid names (e.g., "Bag Kinderarztpraxis") df['is_valid_name'] = df['clean_name'].str.match(r'^[A-Z][a-z]+(?: [A-Z][a-z]+)?$')
Let me break down that regex:
^(?:...): Targets text at the start of the string (non-capturing group so we don't keep the titles)(?:[A-Za-z]+\.)+: Matches stacked abbreviated titles likePr.Doz.Dr.orDr. med.\w+ \w+\s*\([^)]+\): Matches titles with parentheses, likeDr. univ. (Budapest)Herr|Sohn: Matches honorifics directly\s+: Matches the space after the title to remove it too
Step 3: Split Clean Names into First and Last
Now that we have clean names, we can split them into first and last names. We'll use rsplit to split from the right (so if someone has a multi-part first name, it stays intact, and the last word is the surname):
# Split into first name (all words except last) and last name (final word) df[['first_name', 'last_name']] = df['clean_name'].str.rsplit(' ', n=1, expand=True) # Handle single-name cases (like "Heinrich") by filling last_name with the first name df['last_name'] = df['last_name'].fillna(df['first_name'])
Step 4: Check the Results
If you print the output, you'll see clean split names and flags for invalid rows:
print(df[['full_name', 'first_name', 'last_name', 'is_valid_name']])
Sample output (truncated):
full_name first_name last_name is_valid_name 0 Pr.Doz.Dr. Klaus Semmler Facharzt für Frauenhe... Klaus Semmler True 1 Dr. univ. (Budapest) Dalia Lax Dalia Lax True 2 Dr. med. Jovan Stojilkovic Jovan Stojilkovic True ... 7 Sohn Heinrich Heinrich Heinrich True 8 Herr Dr. sc. med. Amadeus Hartwig Amadeus Hartwig True 9 Jasmin Rieche Jasmin Rieche True
Adjustments for Your Full Dataset
- If you run into other titles (like
Frau,Prof., etc.), just add them to theHerr|Sohnpart of the regex. - For rows with trailing extra text (like
Facharzt für Frauenhe...), you can extend the regex to remove text after the name, or filter those rows usingis_valid_name. - If you have multi-part surnames, you might need to tweak the splitting logic, but this works for most standard cases in your sample.
内容的提问来源于stack exchange,提问作者jerof

