Python实现超长字符串递归分割:按指定分隔符自然断行
自定义函数实现自然断行需求
没有现成的Python模块函数能完全匹配你的需求,不过可以通过递归逻辑结合textwrap.wrap()实现自定义函数,以下是完整解决方案:
import textwrap def split_text(text, max_length, separators=('. ', ', ', '? ', '! ')): # 文本长度未超阈值,直接返回 if len(text) <= max_length: return [text] # 在max_length范围内寻找最后一个分隔符的位置 last_sep_pos = -1 for sep in separators: # 扩展查找范围确保能捕获到分隔符本身 pos = text.rfind(sep, 0, max_length + len(sep)) if pos > last_sep_pos: last_sep_pos = pos if last_sep_pos != -1: # 分割出包含分隔符的前半段,递归处理剩余文本 sep_length = len(text[last_sep_pos:last_sep_pos+2]) split_idx = last_sep_pos + sep_length return [text[:split_idx]] + split_text(text[split_idx:], max_length, separators) else: # 无指定分隔符时,用textwrap强制按字符长度分割 return textwrap.wrap(text, width=max_length)
代码说明
- 基础判断:如果当前文本长度不超过设定的
max_length,直接返回包含该文本的列表。 - 分隔符查找:在文本的前
max_length字符范围内,反向遍历所有指定分隔符,找到最后出现的有效分隔符位置,保证断行点最接近阈值且符合自然语义。 - 递归分割:找到有效分隔符后,将文本拆分为前半段(包含分隔符)和剩余部分,对剩余部分重复执行断行逻辑。
- 兜底处理:如果范围内没有找到指定分隔符,调用
textwrap.wrap()按字符长度强制分割。
测试示例
用你提供的测试文本验证:
test_text = "This, is a long, long string. It is around(?) 72 characters! Pretty cool" result = split_text(test_text, 36) print(result)
输出结果:
['This, is a long, long string. ', 'It is around(?) 72 characters! ', 'Pretty cool']
各元素长度分别为30、31、11,均符合不超过36字符的要求。
内容的提问来源于stack exchange,提问作者P A N
相关产品推荐
相关产品推荐

