You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python实现更Pythonic的多余换行符清理功能?

优化多余换行符清理函数

需求背景

将网站HTML转换为纯文本时,往往会产生大量多余的换行符,我们期望最终文本中相邻换行符最多保留1个。原实现代码冗余繁琐,且无法覆盖所有使用场景,需要更Pythonic、更简洁的实现方案。

原实现代码

def clean_up_lines(message_text):
    text_str = str(message_text)
    text_data = text_str.replace(chr(13), "[EOL]")
    text_data = text_data.replace(chr(10), "[EOL]")
    text_data = text_data.replace("\n", "[EOL]")
    text_data = text_data.replace("\r", "[EOL]")
    for x in range(0, 10):
        text_data = text_data.replace("[EOL]       [EOL]", "[EOL]")
        text_data = text_data.replace("[EOL]      [EOL]", "[EOL]")
        text_data = text_data.replace("[EOL]     [EOL]", "[EOL]")
        text_data = text_data.replace("[EOL]    [EOL]", "[EOL]")
        text_data = text_data.replace("[EOL]   [EOL]", "[EOL]")
        text_data = text_data.replace("[EOL]  [EOL]", "[EOL]")
        text_data = text_data.replace("[EOL] [EOL]", "[EOL]")
        text_data = text_data.replace("[EOL][EOL]", "[EOL]")
    for x in range(0, 8):
        text_data = text_data.replace("[EOL][EOL]", "[EOL]")
    text_data = text_data.replace("[EOL]", "\n")
    return text_data

优化后Pythonic实现

直接使用Python内置正则模块即可实现,仅需数行代码,且能覆盖所有场景:

import re

def clean_up_lines(message_text):
    # 匹配连续的换行(含\r、\n)及换行之间的任意空白,替换为单个换行符
    return re.sub(r'[\r\n][\r\n\s]*', '\n', str(message_text))

如果需要同时去掉文本首尾的多余换行,可以再加一步strip处理:

import re

def clean_up_lines(message_text):
    return re.sub(r'[\r\n][\r\n\s]*', '\n', str(message_text)).strip('\n')

实现说明

  • 正则表达式[\r\n][\r\n\s]*会匹配任意回车、换行开头,后续跟任意数量换行、空白字符的连续序列,直接替换为单个\n,完美实现相邻换行最多保留1个的需求
  • 无需自定义占位符和多次循环替换,执行效率远高于原有实现,且不存在空格数量过多时处理失效的问题
  • 兼容所有换行格式(\r、\n、\r\n),覆盖全部场景

内容的提问来源于stack exchange,提问作者krypterro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 12:48:01