如何不使用库按两个及以上换行符分割字符串?
Alright, let's solve this problem without relying on any external libraries. The goal is to split a string whenever there are two or more consecutive newlines, while keeping single newlines intact within each paragraph. Here's a straightforward approach using basic string operations:
Step-by-Step Solution
First, we'll track consecutive newlines as we iterate through the string, replacing any sequence of two or more newlines with a unique separator. Then we can split on that separator and clean up the resulting paragraphs to match your expected output.
Here's the code implementation (in Python, aligned with the list format of your expected output):
def split_on_double_newlines(input_str): processed_chars = [] in_consecutive_newlines = False previous_char = None for char in input_str: if char == '\n': if previous_char == '\n': # Hit a second consecutive newline, ignore further ones until non-newline in_consecutive_newlines = True continue # Single newline, keep it in the paragraph processed_chars.append(char) else: if in_consecutive_newlines: # End of consecutive newlines, add a unique separator processed_chars.append('\x00') # Rare control char to avoid conflicts in_consecutive_newlines = False processed_chars.append(char) previous_char = char # Split using our separator, then clean up each paragraph raw_segments = ''.join(processed_chars).split('\x00') final_paragraphs = [] for segment in raw_segments: # Strip extra newlines/whitespace from edges, skip empty segments cleaned = segment.strip('\n').strip() if cleaned: final_paragraphs.append(cleaned) return final_paragraphs # Test with your input string input_text = """Hello World. I'm very happy today. How are you? Bye.""" print(split_on_double_newlines(input_text)) # Output: ["Hello World.\n I'm very happy today", "How are you?", "Bye"]
How This Works
- Track Consecutive Newlines: We loop through each character, flagging when we hit a second consecutive newline. We ignore any additional newlines until we encounter a non-newline character.
- Mark Split Points: When we exit a sequence of consecutive newlines, we add a unique separator (
\x00, a control character unlikely to appear in regular text) to mark where we should split the string. - Clean Up Paragraphs: After splitting on the separator, we strip extra newlines and whitespace from the edges of each segment, and skip any empty segments that might come from leading/trailing consecutive newlines in the original text.
This method preserves single newlines within paragraphs while using two or more consecutive newlines as split points—exactly what you requested.
内容的提问来源于stack exchange,提问作者FasterSol

