You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

字典按累计单词数≤x拆分键并递归处理新键的实现需求

Solution to Split Dictionary by Cumulative Word Count Limit

Problem Statement

Given a dictionary where each value is a list of strings, we need to split the lists into multiple entries such that the cumulative number of words in each entry does not exceed a specified limit x. When adding the next string would cause the cumulative count to exceed x, we backtrack to the previous string, create a new key with an incrementing suffix (_1, _2, etc.), move the remaining strings to this new key, and repeat the process for the new key until all entries comply with the limit.

Example (x=5)

Input:

initial_dict = {
    'key1': ['abc wer', 'defgh e', 'ij r', 'klmnopqr r', 'e r stuv', 'wxyz aaa', 'extra word'],
    'key2': ['ciao']
}

Output:

output = {
    'key1': ['abc wer', 'defgh e'], 
    'key1_1': ['ij r', 'klmnopqr r'],
    'key1_2': ['e r stuv', 'wxyz aaa'],
    'key1_3': ['extra word'],
    'key2': ['ciao']
}

Approach

  • Count Words per String: Calculate the number of words in each string by splitting on whitespace.
  • Iterate and Split: For each key in the original dictionary:
    • Start with the full list of strings for the key.
    • Track cumulative word count as we iterate through the strings.
    • When adding the next string would exceed the limit x, split the list at the previous position.
    • Assign the first part to the current key (or suffixed key) and move remaining strings to a new suffixed key.
    • Repeat the process for remaining strings until all subsets have cumulative word counts ≤ x.

Solution Code

def count_words(s: str) -> int:
    """Count the number of words in a string (split by whitespace)."""
    return len(s.split())

def split_dict_by_word_limit(initial_dict: dict, x: int) -> dict:
    output = {}
    for key, strings in initial_dict.items():
        current_list = strings.copy()
        suffix = 0
        while current_list:
            # Determine the current key (original or suffixed)
            current_key = key if suffix == 0 else f"{key}_{suffix}"
            total_words = 0
            split_index = 0
            
            # Find the maximum number of strings we can take without exceeding x
            for i, string in enumerate(current_list):
                word_count = count_words(string)
                if total_words + word_count > x:
                    break
                total_words += word_count
                split_index = i + 1
            
            # Assign the subset to the current key
            output[current_key] = current_list[:split_index]
            # Update current_list to the remaining strings
            current_list = current_list[split_index:]
            suffix += 1
    return output

Example Usage

initial_dict = {
    'key1': ['abc wer', 'defgh e', 'ij r', 'klmnopqr r', 'e r stuv', 'wxyz aaa', 'extra word'],
    'key2': ['ciao']
}
x = 5
result = split_dict_by_word_limit(initial_dict, x)
print(result)

Output:

{
    'key1': ['abc wer', 'defgh e'], 
    'key1_1': ['ij r', 'klmnopqr r'],
    'key1_2': ['e r stuv', 'wxyz aaa'],
    'key1_3': ['extra word'],
    'key2': ['ciao']
}

Notes

  • This solution assumes every individual string has a word count ≤ x. If a single string's word count exceeds x, it will be placed in its own key, violating the limit. Add a check for this scenario if needed.
  • The function creates a new dictionary without modifying the original input.

内容的提问来源于stack exchange,提问作者Federicofkt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 14:05:57