You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则提取带贵族头衔的唯一姓名并匹配列表问题求助

Hey there! Let's fix your code step by step to get the exact output you want, then handle the matching with the other_names list.

First, the issue with your current regex is that it only captures titles followed by two words (like Baroness Firstname Surname), but misses valid single-name title entries like Lady Anothername. However, we need to avoid capturing entries like Lady Surname or Lady Firstname which are just parts of another full name. Here's how to do it:

Step 1: Extract Valid Names with Titles

We'll split the process into two parts:

  1. Capture all full title + two-word name combinations.
  2. Capture title + single-word combinations, but exclude any where the single word is a first/last name from the full combinations.

Here's the updated function:

import re

def get_names(text):
    # Pattern for title + two-word names (e.g., Baroness Firstname Surname)
    two_word_pattern = re.compile(r'(Lord|Baroness|Lady|Baron) ([A-Z][a-z]+) ([A-Z][a-z]+)')
    two_word_matches = two_word_pattern.findall(text)
    two_word_names = [' '.join(match) for match in two_word_matches]
    
    # Collect all individual first/last names from the two-word entries to filter out invalid single-word titles
    individual_names = set()
    for title, first, last in two_word_matches:
        individual_names.add(first)
        individual_names.add(last)
    
    # Pattern for title + single-word names (e.g., Lady Anothername)
    one_word_pattern = re.compile(r'(Lord|Baroness|Lady|Baron) ([A-Z][a-z]+)')
    one_word_matches = one_word_pattern.findall(text)
    one_word_candidates = [' '.join(match) for match in one_word_matches]
    
    # Filter out single-word titles that are just parts of full names
    valid_one_word_names = []
    for name in one_word_candidates:
        title, name_part = name.split(maxsplit=1)
        if name_part not in individual_names:
            valid_one_word_names.append(name)
    
    # Combine, remove duplicates, and sort for consistency
    all_valid_names = list(set(two_word_names + valid_one_word_names))
    all_valid_names.sort()
    return all_valid_names

Testing the Function

Let's run it with your sample text:

text = 'Baroness Firstname Surname and Baroness who is also known as Lady Anothername and Lady Surname or Lady Firstname.'
names_lst = get_names(text)
print(names_lst)
# Output: ['Baroness Firstname Surname', 'Lady Anothername']

Perfect, that's exactly the output you wanted!

Step 2: Filter the other_names List

Next, we need to remove any entries in other_names that match the non-title part of our extracted names. Here's how:

other_names = ['Firstname Surname', 'James', 'Simon Smith']

# Get all non-title name parts from our extracted list
non_title_parts = set()
for name in names_lst:
    _, name_part = name.split(maxsplit=1)
    non_title_parts.add(name_part)

# Filter out matching entries
filtered_other_names = [name for name in other_names if name not in non_title_parts]
print(filtered_other_names)
# Output: ['James', 'Simon Smith']

This works because we're checking if each entry in other_names is exactly the full name part (without title) of any extracted noble name, and removing those matches.

内容的提问来源于stack exchange,提问作者aviss

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:48:57