You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将文本拆分为指定数量的完整行(避免单词截断)

问题:如何将字符串拆分为结尾是完整单词的片段?

现有如下Python代码,执行后会按固定字符数拆分字符串,但会把单词从中间截断:

def split(lst, n):
    n = min(n, len(lst))
    k, m = divmod(len(lst), n)
    for i in range(n):
        yield lst[i*k+min(i, m):(i+1)*k+min(i+1, m)] 

siemak = "I look around at my culture, at the ways that people treat each other and at the ways that people treat the land, and I feel that we live in one of those times when language is closed and guarded by gatekeepers. I read Thoreaus Walden and realize that the language we use to speak of the natural world, and of our relationship with the natural world, has changed very little in 150 years. And if the language we use to speak of the natural world is not innovative and engaging then is it any wonder that few young people get excited about nature? I feel that the time has come for language to shine again, to bloom like a flower and lead the way as we begin to speak confidently about the future we want. But this can only happen when we create new words that will serve as vessels for new ideas and new dreams. Long-stable systems and stale old conventions are already breaking down before our eyes, and in the midst of this teetering balance we have an amazing opportunity to rebuild our culture and our relationship with the natural world through language."

porto = list(split(siemak,20))
print(porto)

执行结果(部分片段存在单词截断):

['I look around at my culture, at the ways that people ', 'treat each other and at the ways that people treat th', 'e land, and I feel that we live in one of those times', ...]

需求:调整代码,使拆分后的每个片段结尾为完整单词,或提供其他可行方案。


解决方案

方法一:改进原分割函数,基于空格调整分割点

核心逻辑:先按原思路计算目标分割位置,再向前找到最近的空格(或标点后的空格)作为实际分割点,避免拆分单词。

修改后的代码:

def split_into_full_words(text, n):
    n = min(n, len(text))
    if n == 0:
        return []
    total_length = len(text)
    avg_length = total_length // n
    remainder = total_length % n
    chunks = []
    start = 0
    for i in range(n):
        # 计算目标结束位置,前remainder个片段多分配1个字符
        target_end = start + avg_length + (1 if i < remainder else 0)
        # 非最后一个片段,找最近的空格作为分割点
        if i < n - 1:
            split_pos = text.rfind(' ', start, target_end + 1)
            # 极端情况(无空格)下直接用目标位置
            if split_pos == -1:
                split_pos = target_end
            chunks.append(text[start:split_pos].strip())
            start = split_pos + 1
        else:
            # 最后一个片段直接取到结尾
            chunks.append(text[start:].strip())
    return chunks

siemak = "I look around at my culture, at the ways that people treat each other and at the ways that people treat the land, and I feel that we live in one of those times when language is closed and guarded by gatekeepers. I read Thoreaus Walden and realize that the language we use to speak of the natural world, and of our relationship with the natural world, has changed very little in 150 years. And if the language we use to speak of the natural world is not innovative and engaging then is it any wonder that few young people get excited about nature? I feel that the time has come for language to shine again, to bloom like a flower and lead the way as we begin to speak confidently about the future we want. But this can only happen when we create new words that will serve as vessels for new ideas and new dreams. Long-stable systems and stale old conventions are already breaking down before our eyes, and in the midst of this teetering balance we have an amazing opportunity to rebuild our culture and our relationship with the natural world through language."

porto = split_into_full_words(siemak, 20)
for chunk in porto:
    print(repr(chunk))

方法二:先拆分单词再分组(更直观)

核心逻辑:先把字符串拆成单词列表,再将单词均匀分配到n个组中,最后拼接每个组的单词,确保每个片段都是完整单词的组合。

代码实现:

def split_words_then_group(text, n):
    # 拆分单词(保留带标点的单词,比如"culture,")
    words = text.split()
    if n == 0 or not words:
        return []
    # 计算每组的平均单词数和余数
    avg_words, remainder = divmod(len(words), n)
    chunks = []
    start = 0
    for i in range(n):
        # 前remainder个组多分配1个单词
        end = start + avg_words + (1 if i < remainder else 0)
        chunks.append(' '.join(words[start:end]))
        start = end
    return chunks

siemak = "I look around at my culture, at the ways that people treat each other and at the ways that people treat the land, and I feel that we live in one of those times when language is closed and guarded by gatekeepers. I read Thoreaus Walden and realize that the language we use to speak of the natural world, and of our relationship with the natural world, has changed very little in 150 years. And if the language we use to speak of the natural world is not innovative and engaging then is it any wonder that few young people get excited about nature? I feel that the time has come for language to shine again, to bloom like a flower and lead the way as we begin to speak confidently about the future we want. But this can only happen when we create new words that will serve as vessels for new ideas and new dreams. Long-stable systems and stale old conventions are already breaking down before our eyes, and in the midst of this teetering balance we have an amazing opportunity to rebuild our culture and our relationship with the natural world through language."

porto = split_words_then_group(siemak, 20)
for chunk in porto:
    print(repr(chunk))

内容的提问来源于stack exchange,提问作者AndrewAndrew

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 07:53:13