You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则表达式匹配技术问询:如何匹配文本中的'chapter'与'part'并排除衍生词

Got it, let's tackle this regex problem step by step. Your goal is to match content containing chapter or part, but exclude words where part is just a prefix (like party, particular), while still catching cases like part1, part 1, part: hello world, and standalone part.

First, let's break down why your previous attempts fell short:

  • The regex with [\s,:,0-9]+ fails for standalone part (like in hello world part) because it requires at least one character after part/chapter.
  • Using \bpart\b works for standalone part, but doesn't capture trailing content (like the : 1 in part: 1), and it won't match part1—since digits count as word characters, there's no word boundary between t and 1.

The Solution Regex

Here's a regex that hits all your requirements:

import re
pattern = r'\b(?:chapter|part(?!\p{L}))(?:[\s,:].*)?'

Let's break down each component:

  • \b: Word boundary, ensures we're matching part/chapter as distinct terms (no extra leading letters, like in xpart).
  • (?:chapter|part(?!\p{L})): A non-capturing group to target valid terms:
    • chapter: No restrictions here since you didn't mention excluding its derivatives.
    • part(?!\p{L}): Matches part only if it's not followed by a letter (\p{L} matches any Unicode letter). This directly excludes derivatives like party, particular, and participant.
  • (?:[\s,:].*)?: Optional non-capturing group to grab trailing content:
    • [\s,:]: Matches common separators (space, colon, comma)—adjust this set if you need to include other characters like hyphens.
    • .*: Matches all remaining content until the end of the line (use .*? instead for non-greedy matching if you have multiple matches in one string).
    • ?: Makes this entire group optional, so standalone part/chapter are still matched.

Testing Against Your Samples

Let's verify with your test strings:

  • s1='part1': Matches 'part1'
  • s2='part 1': Matches 'part 1'
  • s3='part: 1': Matches 'part: 1'
  • s4='part hello world': Matches 'part hello world'
  • s5='part: hello world': Matches 'part: hello world'
  • s6='hello world part': Matches 'part'

And it correctly ignores derivatives like party, particular, and participant.

内容的提问来源于stack exchange,提问作者Cihan Senol

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 11:27:38