如何从长文本中用正则提取`says to`后到时间戳前的指定内容
正确实现代码
首先清理输入中的HTML标签,再通过正则匹配目标内容:
import re from html import unescape # 清除字符串中的HTML标签 def remove_html_tags(text): clean_rule = re.compile('<.*?>') return re.sub(clean_rule, '', text) # 预处理原始字符串 cleaned_string = remove_html_tags(unescape(string)) # 正则匹配目标内容 match_pattern = r'says to\s*(.*?)(?=\s*\(\d{4}-\d{2}-\d{2}|\s*---|\s*\*\*\*|$)' text = re.findall(match_pattern, cleaned_string, flags=re.DOTALL) # 清理每条结果的多余换行、空白 text = [re.sub(r'\s+', ' ', item.strip()) for item in text if item.strip()]
运行后得到的text数组即可符合需求。
原写法错误说明
- 存在语法错误:正则末尾的左括号没有配对,运行会直接抛出正则语法异常
- 匹配规则错误:
\D只会匹配非数字字符,但目标内容包含-5、0等数字,会直接截断匹配结果 - 缺少终止匹配逻辑:没有设置到下一个时间戳就停止匹配的规则,也没有适配跨行场景,无法正确提取跨多行的内容
内容的提问来源于stack exchange,提问作者new_learner
相关产品推荐
相关产品推荐

