You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中为段落句子设置单词限制并拆分存入列表?

文本拆分实现方案

需求说明

给定如下英文文本代码:

sent = 'Python is dynamically-typed and garbage-collected. It supports multiple programming paradigms, including structured (particularly procedural), object-oriented and functional programming.'

需要将该文本拆分为每个片段最多包含5个单词的元素,并存入列表,预期输出:

sent_list = ['Python is dynamically-typed and garbage-collected.', 'It supports multiple programming paradigms,', 'including structured (particularly procedural), object-oriented', 'and functional programming.']

中文翻译后的需求

给定如下中文文本代码:

sent = 'Python是动态类型且支持垃圾回收的语言。它支持多种编程范式,包括结构化(尤其是过程式)、面向对象和函数式编程。'

需要将该文本拆分为每个片段最多包含5个词的元素,并将这些元素存入列表,预期输出如下:

sent_list = ['Python是动态类型且支持垃圾回收的语言。', '它支持多种编程范式,', '包括结构化(尤其是过程式)、面向对象', '和函数式编程。']

实现代码

def split_sentence(sent, max_words=5):
    # 拆分单词/词,中文若需精准分词可替换为jieba等库的分词逻辑
    words = sent.split()
    result = []
    current_group = []
    for word in words:
        current_group.append(word)
        if len(current_group) == max_words:
            result.append(' '.join(current_group))
            current_group = []
    # 处理剩余未凑满数量的词
    if current_group:
        result.append(' '.join(current_group))
    return result

# 处理英文原句
sent_en = 'Python is dynamically-typed and garbage-collected. It supports multiple programming paradigms, including structured (particularly procedural), object-oriented and functional programming.'
sent_list_en = split_sentence(sent_en)
print(f"sent_list = {sent_list_en}")

# 处理中文翻译句
sent_zh = 'Python是动态类型且支持垃圾回收的语言。它支持多种编程范式,包括结构化(尤其是过程式)、面向对象和函数式编程。'
sent_list_zh = split_sentence(sent_zh)
print(f"sent_list = {sent_list_zh}")

说明

  • 核心逻辑是遍历拆分后的词列表,每累计到指定数量的词就组合成一个片段存入结果列表,最后处理剩余未凑满数量的词。
  • 中文若需要更精准的分词效果,可引入jieba库替换split()方法,实现按语义分词后再进行分组。

内容的提问来源于stack exchange,提问作者waji

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 01:20:31