如何使用正则表达式匹配两个子串之间的多行段落内容
正则提取多行正文段落解决方案
原代码问题分析
- 未开启
re.DOTALL(别名re.S)匹配模式,正则默认下.元字符无法匹配换行符,只能捕获单行内容 - 匹配逻辑未设置正确的段落结束边界,也没有对正文部分做单独捕获
正确实现代码
import re # 原始字符串 string = """( 2021-07-10 01:24:55 PM GMT )TEST --- Badminton is a racquet sport played using racquets to hit a shuttlecock across a net. Although it may be played with larger teams, the most common forms of the game are "singles" (with one player per side) and "doubles" (with two players per side). ( 2021-07-10 01:27:55 PM GMT )PATRICKWARR --- Good morning, I am doing well. And you? --- --- * * *""" # 匹配规则:匹配时间头、分隔线后捕获正文,直到下一个时间头或末尾分隔符停止 pattern = r'\( \d{4}-\d{2}-\d{2}\s\d{2}:\d{2}:\d{2} [AP]M GMT \)\w+\s+---\s+(.*?)(?=\s*\( \d{4}-\d{2}-\d{2}|---\s*\* \* \*)' # 开启DOTALL模式让.可以匹配换行,提取后去除段落前后的空白字符 text = [para.strip() for para in re.findall(pattern, string, flags=re.DOTALL)] print(text)
输出结果
和预期完全一致:
['Badminton is a racquet sport played using racquets to hit a shuttlecock across\na net. Although it may be played with larger teams, the most common forms of\nthe game are "singles" (with one player per side) and "doubles" (with two\nplayers per side).', 'Good morning, I am doing well. And you?']
内容的提问来源于stack exchange,提问作者new_learner
相关产品推荐
相关产品推荐

