如何将列表首子列表字符串拆分填充至多子列表且不拆分单词?
问题需求
给定嵌套列表格式如下:
[["Hi, this is Tesa form the sales departmet. Iam working here from"],[""],[""]]
需将首子列表的字符串拆分后填充到所有子列表中,满足:
- 子列表数量可为任意n,除首子列表外其余均为空
- 拆分时不得切割单词,无需保留语句语义
示例预期输出:
[["Hi, this is Tesa."],["form the sales departmet"],["Iam working here from"]]
现有实现问题
当前代码通过指定字符位置切割字符串,会出现拆分单词的情况,输出如下:
[['Hi, this is Tes'], ['a form the sales departm'], ['et. Iam working here from']]
且数据集规模较大,需要更高效简便的解决方案。
解决方案
以下实现通过先拆分单词再均分的方式,彻底避免拆分单词,同时适配任意数量的子列表:
def explode_string(array): original_text = array[0][0] num_sublists = len(array) # 拆分文本为单词列表 words = original_text.split() # 计算每个子列表分配的单词数,尽量均分 base_count = len(words) // num_sublists extra = len(words) % num_sublists result = [] current_idx = 0 for i in range(num_sublists): # 前extra个分组多分配一个单词 end_idx = current_idx + base_count + (1 if i < extra else 0) # 拼接单词为字符串并加入结果 result.append([' '.join(words[current_idx:end_idx])]) current_idx = end_idx return result # 测试示例 original_array = [ ["Hi, this is Tesa form the sales departmet. Iam working here from"], [""], [""] ] print(explode_string(original_array))
输出结果:
[['Hi, this is Tesa'], ['form the sales departmet.'], ['Iam working here from']]
该方案优势:
- 完全规避单词拆分问题,符合需求
- 无需手动指定拆分位置,自动适配任意子列表数量
- 时间复杂度为O(m)(m为单词总数),处理大规模数据集效率更高
内容的提问来源于stack exchange,提问作者Bhargav
相关产品推荐
相关产品推荐

