You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将列表首子列表字符串拆分填充至多子列表且不拆分单词?

问题需求

给定嵌套列表格式如下:

[["Hi, this is Tesa form the sales departmet. Iam working here from"],[""],[""]]

需将首子列表的字符串拆分后填充到所有子列表中,满足:

  • 子列表数量可为任意n,除首子列表外其余均为空
  • 拆分时不得切割单词,无需保留语句语义

示例预期输出:

[["Hi, this is Tesa."],["form the sales departmet"],["Iam working here from"]]
现有实现问题

当前代码通过指定字符位置切割字符串,会出现拆分单词的情况,输出如下:

[['Hi, this is Tes'], ['a form the sales departm'], ['et. Iam working here from']]

且数据集规模较大,需要更高效简便的解决方案。

解决方案

以下实现通过先拆分单词再均分的方式,彻底避免拆分单词,同时适配任意数量的子列表:

def explode_string(array):
    original_text = array[0][0]
    num_sublists = len(array)
    # 拆分文本为单词列表
    words = original_text.split()
    # 计算每个子列表分配的单词数,尽量均分
    base_count = len(words) // num_sublists
    extra = len(words) % num_sublists
    
    result = []
    current_idx = 0
    for i in range(num_sublists):
        # 前extra个分组多分配一个单词
        end_idx = current_idx + base_count + (1 if i < extra else 0)
        # 拼接单词为字符串并加入结果
        result.append([' '.join(words[current_idx:end_idx])])
        current_idx = end_idx
    
    return result

# 测试示例
original_array = [
    ["Hi, this is Tesa form the sales departmet. Iam working here from"],
    [""],
    [""]
]

print(explode_string(original_array))

输出结果:

[['Hi, this is Tesa'], ['form the sales departmet.'], ['Iam working here from']]

该方案优势:

  • 完全规避单词拆分问题,符合需求
  • 无需手动指定拆分位置,自动适配任意子列表数量
  • 时间复杂度为O(m)(m为单词总数),处理大规模数据集效率更高

内容的提问来源于stack exchange,提问作者Bhargav

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 02:49:53