如何将字符串分割为包含完整单词的等长子串(不使用split()方法)
Solution to Split Text into Equal Chunks with Complete Words
I get it—your current approach splits text strictly by length, which cuts words right in the middle, and you don’t want to pre-split words because each chunk needs to be a full string for other processes to handle. Here’s a way to adjust your code to find natural word boundaries (spaces) near your target split points, so every chunk ends with a complete word:
Step-by-Step Explanation
Instead of splitting exactly at the calculated length position, we:
- Calculate the approximate size each chunk should be
- For each chunk, find the next space after the target end position to avoid cutting words
- Handle the last chunk separately to grab the remaining text without overflow
Updated Code
def split_text_into_full_word_chunks(text, n_divisions): chunks = [] total_length = len(text) approx_chunk_size = total_length // n_divisions current_start = 0 for i in range(n_divisions): if i == n_divisions - 1: # Last chunk: take everything from current start to end chunks.append(text[current_start:]) break target_end = current_start + approx_chunk_size # Find the next space after the target end to avoid cutting words next_space = text.find(' ', target_end) # Fallback: if no space found (unlikely), take the rest of the text if next_space == -1: chunks.append(text[current_start:]) break # Add the chunk with complete words, move start to after the space chunks.append(text[current_start:next_space]) current_start = next_space + 1 return chunks def main(): text = "Lorem Ipsum is simply dummy text of the printing and " \ "typesetting industry. Lorem Ipsum has been the industry's " \ "standard dummy text ever since the 1500s, when an unknown " \ "printer took a galley of type and scrambled it to make a type " \ "specimen book. It has survived not only five centuries, but also " \ "the leap into electronic typesetting, remaining essentially unchanged. " \ "It was popularised in the 1960s with the release of Letraset sheets " \ "containing Lorem Ipsum passages, and more recently with desktop publishing " \ "software like Aldus PageMaker including versions of Lorem Ipsum" n_divisions = 5 chunks = split_text_into_full_word_chunks(text, n_divisions) for i, chunk in enumerate(chunks): print(f"{i} Division: {chunk}") # Optional check to confirm no truncated words print(f"Chunk ends with complete word: {chunk[-1] not in 'abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ' or text[current_start] == ' '}\n") if __name__ == '__main__': main()
How It Works
- Approximate Chunk Size: We start with
total_length // n_divisionsto get a rough idea of how long each chunk should be, keeping splits as balanced as possible. - Word Boundary Detection: Using
text.find(' ', target_end)we look for the next space after our target split point. This extends the chunk to the end of the current word, avoiding truncation. - Last Chunk Handling: The final chunk takes whatever is left of the text, so we don’t have to worry about miscalculations for the last piece.
Key Notes
- This works seamlessly with punctuation (commas, periods) since they’re typically followed by spaces in standard prose.
- We avoid using
split()entirely, so each chunk remains a full, unprocessed string ready for your other processes to handle.
内容的提问来源于stack exchange,提问作者Chariot
相关产品推荐
相关产品推荐

