You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Pandas入门:如何从含头衔的14k行数据集中拆分姓名

Splitting First and Last Names from a Dataset with Titles (Python/Pandas)

Hey there! Let's break down how to split first and last names from your messy 14k-row dataset—all those titles, honorifics, and extra text can be tricky, but with regex and Pandas' string tools, we can clean this up effectively.

First, let's look at your data patterns: you've got everything from stacked academic titles (Pr.Doz.Dr.), titles with parentheses (Dr. univ. (Budapest)), honorifics like Herr/Sohn, and even plain names with no titles. Our plan is to strip out all the leading title/honorific text first, then split the remaining clean name into first and last parts.

Step 1: Load Your Data

First, let's get your data into a Pandas DataFrame. I'll use your sample data for this example:

import pandas as pd

# Sample data matching your example
sample_data = {
    'full_name': [
        'Pr.Doz.Dr. Klaus Semmler Facharzt für Frauenhe...',
        'Dr. univ. (Budapest) Dalia Lax',
        'Dr. med. Jovan Stojilkovic',
        'Dr. med. Dirk Schneider',
        'Marc Scheuermann',
        'Bag Kinderarztpraxis',
        'Herr Ulrich Bromig',
        'Sohn Heinrich',
        'Herr Dr. sc. med. Amadeus Hartwig',
        'Jasmin Rieche'
    ]
}
df = pd.DataFrame(sample_data)

Step 2: Strip Leading Titles/Honorifics

We'll use a regular expression to target and remove all the leading title patterns. Here's a regex that covers all the cases in your sample:

# Regex pattern to match leading titles/honorifics
title_regex = r'^(?:(?:[A-Za-z]+\.)+|\w+ \w+\s*\([^)]+\)|Herr|Sohn)\s+'

# Remove titles to get a clean name string
df['clean_name'] = df['full_name'].str.replace(title_regex, '', regex=True)

# Flag rows that don't look like valid names (e.g., "Bag Kinderarztpraxis")
df['is_valid_name'] = df['clean_name'].str.match(r'^[A-Z][a-z]+(?: [A-Z][a-z]+)?$')

Let me break down that regex:

  • ^(?:...): Targets text at the start of the string (non-capturing group so we don't keep the titles)
  • (?:[A-Za-z]+\.)+: Matches stacked abbreviated titles like Pr.Doz.Dr. or Dr. med.
  • \w+ \w+\s*\([^)]+\): Matches titles with parentheses, like Dr. univ. (Budapest)
  • Herr|Sohn: Matches honorifics directly
  • \s+: Matches the space after the title to remove it too

Step 3: Split Clean Names into First and Last

Now that we have clean names, we can split them into first and last names. We'll use rsplit to split from the right (so if someone has a multi-part first name, it stays intact, and the last word is the surname):

# Split into first name (all words except last) and last name (final word)
df[['first_name', 'last_name']] = df['clean_name'].str.rsplit(' ', n=1, expand=True)

# Handle single-name cases (like "Heinrich") by filling last_name with the first name
df['last_name'] = df['last_name'].fillna(df['first_name'])

Step 4: Check the Results

If you print the output, you'll see clean split names and flags for invalid rows:

print(df[['full_name', 'first_name', 'last_name', 'is_valid_name']])

Sample output (truncated):

full_name first_name   last_name  is_valid_name
0  Pr.Doz.Dr. Klaus Semmler Facharzt für Frauenhe...       Klaus     Semmler           True
1               Dr. univ. (Budapest) Dalia Lax       Dalia         Lax           True
2                    Dr. med. Jovan Stojilkovic       Jovan  Stojilkovic           True
...
7                                Sohn Heinrich    Heinrich    Heinrich           True
8        Herr Dr. sc. med. Amadeus Hartwig       Amadeus     Hartwig           True
9                              Jasmin Rieche      Jasmin      Rieche           True

Adjustments for Your Full Dataset

  • If you run into other titles (like Frau, Prof., etc.), just add them to the Herr|Sohn part of the regex.
  • For rows with trailing extra text (like Facharzt für Frauenhe...), you can extend the regex to remove text after the name, or filter those rows using is_valid_name.
  • If you have multi-part surnames, you might need to tweak the splitting logic, but this works for most standard cases in your sample.

内容的提问来源于stack exchange,提问作者jerof

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 17:53:00