You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则表达式匹配多个姓名:中间Lookbehind无法正常工作

Extracting 1-3 Names from "By..." Strings: A Reliable Regex Solution

I see you're struggling to reliably extract two-word names from strings formatted like "By Person One, Person Two and Person Three"—your current regex has issues with lookbehind assertions breaking matches for different separator combinations. Let's fix this with a cleaner approach that handles all your example cases.

The Core Problem

Your original regex overcomplicates things with lookbehinds that create hard dependencies on specific separators. Instead, we can either capture names directly with flexible separator handling, or split the problem into two simple steps for even more reliability.

Option 1: Single Regex with Capture Groups (For 1-3 Names)

This regex will capture each name as a separate group, handling both comma and "and" separators correctly. It works for all your test cases:

(?<=By\s)(?:(\w+\s\w+)(?:,\s*)?(?:and\s)?(\w+\s\w+)?(?:(?:,\s*)?and\s)?(\w+\s\w+)?)
  • The (?<=By\s) positive lookbehind ensures we only match text right after "By ".
  • Each name is captured as (\w+\s\w+) (two words separated by space).
  • Optional groups (?:,\s*)? and (?:and\s)? handle the different separator combinations (comma, "and", or comma + "and").

After matching, you'll get up to three capture groups—just ignore any empty groups if there are fewer than three names.

Option 2: Global Match (Simpler, For Any Number of Names)

If your regex engine supports global matching (like Python's re.findall, JavaScript's g flag), this is even easier:

  1. First, extract the substring after "By " using:
    (?<=By\s).+
    
  2. Then, use this regex globally to pull out every two-word name:
    \w+\s\w+
    

Example in Python:

import re

sample_text = "By Person Two and Person Three"
# Get the part after "By "
names_section = re.search(r'(?<=By\s).+', sample_text).group()
# Extract all names
names_list = re.findall(r'\w+\s\w+', names_section)
# Result: ["Person Two", "Person Three"]

Why Your Original Regex Broke

The (?<=and ) lookbehind in your pattern creates a requirement that "and " must immediately precede the third name. But in cases like "Person Two and Person Three", your regex still expects an optional comma before that "and"—this mismatch causes the entire match to fail. Lookbehinds work best when they're not tied to optional groups, as they need the preceding text to exist even if the group is marked as optional.

Edge Case Notes

  • If names might include middle initials or more than two words, adjust the name pattern to something like [\w\.]+\s[\w\.]+\s?[\w\.]* (but this depends on your specific name format).
  • Add the i (case-insensitive) flag to handle lowercase "by" in input text.

内容的提问来源于stack exchange,提问作者Collin Chaffin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:11:09