寻求构建Python正则表达式匹配特定句子模式的解决方案
Python正则表达式匹配特定格式内容的解决方案
先看需求示例:
pattern1 = '.She is greatThis is annoyingWhy u do this' pattern2 = '.Weirdly specificThis sentence is longer than the other oneSee this is great' example = 'He went such dare good mr fact. The small own seven saved man age no offer. Suspicion did mrs nor furniture smallness. Scale whole downs often leave not eat. An expression reasonably cultivated indulgence mr he surrounded instrument. Gentleman eat and consisted are pronounce distrusts.This is where the fun startsSummer is really bothersome this yearShe is out of ideas' example_pattern_goal = 'This is where the fun startsSummer is really bothersome this yearShe is out of ideas'
核心匹配规则
- 目标内容以点号(.)开头,整体不含数字
- 点号后紧跟首字母大写的单词
- 后续内容中至少包含两处单词内部出现大写字母的情况(比如示例里
Summer、She这类大写开头的拼接式单词)
原正则的问题
你之前写的正则'.\b[A-Z]\w+[\s\w]+\b\w+[A-Z]\w+\b[\s\w]+\b\w+[A-Z]\w+\b[\s\w]+'有几个关键缺陷:
- 点号
.是正则通配符,未转义会匹配任意字符,无法精准匹配字面点号 \b位置错误,点号后直接接大写字母,此处不需要单词边界匹配- 逻辑僵化,硬编码两次大写位置,无法灵活满足“至少两处”的要求,也没处理“不含数字”的限制
正确的正则表达式实现
以下是符合要求的正则及Python测试代码:
import re # 正则表达式 regex = r'\.(?=[A-Z][^\d]+(?:[a-z][A-Z][^\d]+){2,})[A-Z][^\d]+' # 测试示例文本 example = 'He went such dare good mr fact. The small own seven saved man age no offer. Suspicion did mrs nor furniture smallness. Scale whole downs often leave not eat. An expression reasonably cultivated indulgence mr he surrounded instrument. Gentleman eat and consisted are pronounce distrusts.This is where the fun startsSummer is really bothersome this yearShe is out of ideas' match = re.search(regex, example) if match: # 去除开头的点号,得到目标内容 result = match.group()[1:] print(result) # 输出:This is where the fun startsSummer is really bothersome this yearShe is out of ideas
正则各部分解释
\.:转义后匹配字面意义的点号(?=[A-Z][^\d]+(?:[a-z][A-Z][^\d]+){2,}):正向预查,确保后续内容符合规则:[A-Z][^\d]+:匹配点号后首字母大写的单词,且后续内容不含数字(?:[a-z][A-Z][^\d]+){2,}:匹配至少两次“小写字母后接大写字母”的场景(即满足至少两处大写开头的拼接单词),(?:...)为非捕获组,{2,}表示至少匹配两次
[A-Z][^\d]+:实际匹配点号后的目标内容,从大写字母开始,直到遇到数字(因规则规定不含数字,会匹配完整的目标片段)
验证其他示例
对pattern1和pattern2测试:
pattern1 = '.She is greatThis is annoyingWhy u do this' match1 = re.search(regex, pattern1) print(match1.group()[1:]) # 输出:She is greatThis is annoyingWhy u do this pattern2 = '.Weirdly specificThis sentence is longer than the other oneSee this is great' match2 = re.search(regex, pattern2) print(match2.group()[1:]) # 输出:Weirdly specificThis sentence is longer than the other oneSee this is great
内容的提问来源于stack exchange,提问作者blue-create
相关产品推荐
相关产品推荐

