You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现超长字符串递归分割:按指定分隔符自然断行

自定义函数实现自然断行需求

没有现成的Python模块函数能完全匹配你的需求,不过可以通过递归逻辑结合textwrap.wrap()实现自定义函数,以下是完整解决方案:

import textwrap

def split_text(text, max_length, separators=('. ', ', ', '? ', '! ')):
    # 文本长度未超阈值,直接返回
    if len(text) <= max_length:
        return [text]
    
    # 在max_length范围内寻找最后一个分隔符的位置
    last_sep_pos = -1
    for sep in separators:
        # 扩展查找范围确保能捕获到分隔符本身
        pos = text.rfind(sep, 0, max_length + len(sep))
        if pos > last_sep_pos:
            last_sep_pos = pos
    
    if last_sep_pos != -1:
        # 分割出包含分隔符的前半段,递归处理剩余文本
        sep_length = len(text[last_sep_pos:last_sep_pos+2])
        split_idx = last_sep_pos + sep_length
        return [text[:split_idx]] + split_text(text[split_idx:], max_length, separators)
    else:
        # 无指定分隔符时,用textwrap强制按字符长度分割
        return textwrap.wrap(text, width=max_length)

代码说明

  1. 基础判断:如果当前文本长度不超过设定的max_length,直接返回包含该文本的列表。
  2. 分隔符查找:在文本的前max_length字符范围内,反向遍历所有指定分隔符,找到最后出现的有效分隔符位置,保证断行点最接近阈值且符合自然语义。
  3. 递归分割:找到有效分隔符后,将文本拆分为前半段(包含分隔符)和剩余部分,对剩余部分重复执行断行逻辑。
  4. 兜底处理:如果范围内没有找到指定分隔符,调用textwrap.wrap()按字符长度强制分割。

测试示例

用你提供的测试文本验证:

test_text = "This, is a long, long string. It is around(?) 72 characters! Pretty cool"
result = split_text(test_text, 36)
print(result)

输出结果:

['This, is a long, long string. ', 'It is around(?) 72 characters! ', 'Pretty cool']

各元素长度分别为30、31、11,均符合不超过36字符的要求。

内容的提问来源于stack exchange,提问作者P A N

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 03:41:30