如何在Python中为段落句子设置单词限制并拆分存入列表?
文本拆分实现方案
需求说明
给定如下英文文本代码:
sent = 'Python is dynamically-typed and garbage-collected. It supports multiple programming paradigms, including structured (particularly procedural), object-oriented and functional programming.'
需要将该文本拆分为每个片段最多包含5个单词的元素,并存入列表,预期输出:
sent_list = ['Python is dynamically-typed and garbage-collected.', 'It supports multiple programming paradigms,', 'including structured (particularly procedural), object-oriented', 'and functional programming.']
中文翻译后的需求
给定如下中文文本代码:
sent = 'Python是动态类型且支持垃圾回收的语言。它支持多种编程范式,包括结构化(尤其是过程式)、面向对象和函数式编程。'
需要将该文本拆分为每个片段最多包含5个词的元素,并将这些元素存入列表,预期输出如下:
sent_list = ['Python是动态类型且支持垃圾回收的语言。', '它支持多种编程范式,', '包括结构化(尤其是过程式)、面向对象', '和函数式编程。']
实现代码
def split_sentence(sent, max_words=5): # 拆分单词/词,中文若需精准分词可替换为jieba等库的分词逻辑 words = sent.split() result = [] current_group = [] for word in words: current_group.append(word) if len(current_group) == max_words: result.append(' '.join(current_group)) current_group = [] # 处理剩余未凑满数量的词 if current_group: result.append(' '.join(current_group)) return result # 处理英文原句 sent_en = 'Python is dynamically-typed and garbage-collected. It supports multiple programming paradigms, including structured (particularly procedural), object-oriented and functional programming.' sent_list_en = split_sentence(sent_en) print(f"sent_list = {sent_list_en}") # 处理中文翻译句 sent_zh = 'Python是动态类型且支持垃圾回收的语言。它支持多种编程范式,包括结构化(尤其是过程式)、面向对象和函数式编程。' sent_list_zh = split_sentence(sent_zh) print(f"sent_list = {sent_list_zh}")
说明
- 核心逻辑是遍历拆分后的词列表,每累计到指定数量的词就组合成一个片段存入结果列表,最后处理剩余未凑满数量的词。
- 中文若需要更精准的分词效果,可引入
jieba库替换split()方法,实现按语义分词后再进行分组。
内容的提问来源于stack exchange,提问作者waji
相关产品推荐
相关产品推荐

