You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将字符串分割为包含完整单词的等长子串(不使用split()方法)

Solution to Split Text into Equal Chunks with Complete Words

I get it—your current approach splits text strictly by length, which cuts words right in the middle, and you don’t want to pre-split words because each chunk needs to be a full string for other processes to handle. Here’s a way to adjust your code to find natural word boundaries (spaces) near your target split points, so every chunk ends with a complete word:

Step-by-Step Explanation

Instead of splitting exactly at the calculated length position, we:

  • Calculate the approximate size each chunk should be
  • For each chunk, find the next space after the target end position to avoid cutting words
  • Handle the last chunk separately to grab the remaining text without overflow

Updated Code

def split_text_into_full_word_chunks(text, n_divisions):
    chunks = []
    total_length = len(text)
    approx_chunk_size = total_length // n_divisions
    current_start = 0
    
    for i in range(n_divisions):
        if i == n_divisions - 1:
            # Last chunk: take everything from current start to end
            chunks.append(text[current_start:])
            break
        
        target_end = current_start + approx_chunk_size
        # Find the next space after the target end to avoid cutting words
        next_space = text.find(' ', target_end)
        
        # Fallback: if no space found (unlikely), take the rest of the text
        if next_space == -1:
            chunks.append(text[current_start:])
            break
        
        # Add the chunk with complete words, move start to after the space
        chunks.append(text[current_start:next_space])
        current_start = next_space + 1
    
    return chunks

def main():
    text = "Lorem Ipsum is simply dummy text of the printing and " \
           "typesetting industry. Lorem Ipsum has been the industry's " \
           "standard dummy text ever since the 1500s, when an unknown " \
           "printer took a galley of type and scrambled it to make a type " \
           "specimen book. It has survived not only five centuries, but also " \
           "the leap into electronic typesetting, remaining essentially unchanged. " \
           "It was popularised in the 1960s with the release of Letraset sheets " \
           "containing Lorem Ipsum passages, and more recently with desktop publishing " \
           "software like Aldus PageMaker including versions of Lorem Ipsum"
    n_divisions = 5
    chunks = split_text_into_full_word_chunks(text, n_divisions)
    
    for i, chunk in enumerate(chunks):
        print(f"{i} Division: {chunk}")
        # Optional check to confirm no truncated words
        print(f"Chunk ends with complete word: {chunk[-1] not in 'abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ' or text[current_start] == ' '}\n")

if __name__ == '__main__':
    main()

How It Works

  • Approximate Chunk Size: We start with total_length // n_divisions to get a rough idea of how long each chunk should be, keeping splits as balanced as possible.
  • Word Boundary Detection: Using text.find(' ', target_end) we look for the next space after our target split point. This extends the chunk to the end of the current word, avoiding truncation.
  • Last Chunk Handling: The final chunk takes whatever is left of the text, so we don’t have to worry about miscalculations for the last piece.

Key Notes

  • This works seamlessly with punctuation (commas, periods) since they’re typically followed by spaces in standard prose.
  • We avoid using split() entirely, so each chunk remains a full, unprocessed string ready for your other processes to handle.

内容的提问来源于stack exchange,提问作者Chariot

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 12:02:46