如何用正则表达式合并拆分多行的单行文本并保留独立句子
正则表达式解决方案
核心替换规则
使用正则匹配非空行末尾的换行符+后续小写字母开头的内容,将其替换为空格连接的形式,同时保留空行和首字母大写的独立句子。
- 正则模式:
(?<!\n)\n\s*([a-z]) - 替换字符串:
$1
正则逻辑解释
(?<!\n):负向后断言,确保当前换行符的前一个字符不是换行(排除空行场景,避免合并空行)\n:匹配需要处理的换行符\s*:匹配换行后可能存在的任意空白字符(空格、制表符等),统一替换为单个空格([a-z]):捕获下一行开头的小写字母,替换时保留该字母,保证文本连贯性
代码示例(Python)
import re original_str = '''I love to eat apple and bananas but at the same time, I do not like to eat oranges and I am not a fan of vegetables. I also like to eat biscuits. I also like to eat crackers. Meanwhile, oddly enough, I have never liked ice creams or chocolates because I am never a sweet-toothed.''' # 执行替换 new_str = re.sub(r'(?<!\n)\n\s*([a-z])', r' \1', original_str) print(new_str)
输出结果
I love to eat apple and bananas but at the same time, I do not like to eat oranges and I am not a fan of vegetables. I also like to eat biscuits. I also like to eat crackers. Meanwhile, oddly enough, I have never liked ice creams or chocolates because I am never a sweet-toothed.
内容的提问来源于stack exchange,提问作者Noor Ameera Anas
相关产品推荐
相关产品推荐

