C#中如何高效逻辑拆分长字符串至指定长度字段?
Great question! Your current code handles sentence-level splitting well, but it misses critical edge cases where a single sentence (or even a single word) exceeds your field length limit. Let's walk through the optimal approach, which builds on your existing logic while adding support for word and character-level splitting when needed.
Why Your Current Approach Needs Adjustment
Your code works fine when all sentences fit within the field length, but it will fail silently when a sentence is longer than splitLength—it'll just dump the entire sentence into a new field, violating your length constraint. We need nested handling for long sentences: first split into words, then split individual words into characters if they're still too long.
The Optimal Strategy: Layered Splitting
The most efficient way to implement this is hierarchical processing:
- First split the entire text into sentences (using your existing regex logic—this is optimal because sentences are the highest-priority split point).
- For each sentence, check if it fits into the current field's remaining space.
- If it doesn't, first flush the current field (with
[cont]), then process the sentence by splitting into words. - For each word, check if it fits into the remaining space. If not, flush the field and handle the word:
- If the word is shorter than the field limit, add it to the new field.
- If the word is longer than the field limit, split it character-by-character into chunks of the maximum allowed length.
This approach is efficient because it minimizes repeated scans of the input text and leverages StringBuilder to avoid costly string concatenations.
Complete Implementation
Here's the enhanced code that covers all your requirements:
using System.Text; using System.Text.RegularExpressions; public static IEnumerable<string> SplitTextToFields(string workText, int maximumLengthOfField) { const string splitIndicator = " [cont]"; int splitLength = maximumLengthOfField - splitIndicator.Length; // Edge case: if field length is smaller than the indicator, throw or handle gracefully if (splitLength <= 0) throw new ArgumentException("Field length must be larger than the length of the continuation indicator."); // Split text into sentences (preserves punctuation at the end of each sentence) string[] sentences = Regex.Split(workText, @"(?<=\.!\?)\s+"); StringBuilder currentLine = new StringBuilder(); foreach (string sentence in sentences) { // Calculate length needed to add this sentence to the current line (including space if needed) int requiredLength = currentLine.Length + (currentLine.Length > 0 ? 1 : 0) + sentence.Length; if (requiredLength <= splitLength) { // Sentence fits: add to current line with space if needed if (currentLine.Length > 0) currentLine.Append(' '); currentLine.Append(sentence); } else { // Flush current line if it's not empty if (currentLine.Length > 0) { yield return currentLine.ToString() + splitIndicator; currentLine.Clear(); } // Now process the long sentence by splitting into words string[] words = sentence.Split(new[] { ' ' }, StringSplitOptions.RemoveEmptyEntries); foreach (string word in words) { int wordRequiredLength = currentLine.Length + (currentLine.Length > 0 ? 1 : 0) + word.Length; if (wordRequiredLength <= splitLength) { // Word fits: add to current line if (currentLine.Length > 0) currentLine.Append(' '); currentLine.Append(word); } else { // Flush current line if not empty if (currentLine.Length > 0) { yield return currentLine.ToString() + splitIndicator; currentLine.Clear(); } // Handle word that's longer than splitLength: split into characters if (word.Length > splitLength) { int index = 0; while (index < word.Length) { int chunkLength = Math.Min(splitLength, word.Length - index); string chunk = word.Substring(index, chunkLength); // If it's not the last chunk, add the continuation indicator if (index + chunkLength < word.Length) { yield return chunk + splitIndicator; } else { currentLine.Append(chunk); } index += chunkLength; } } else { // Word fits in empty line: add it currentLine.Append(word); } } } } } // Yield the final line (no continuation indicator) if (currentLine.Length > 0) yield return currentLine.ToString(); }
Key Optimizations & Notes
- Regex Sentence Splitting: Your original regex
(?<=\.!\?)\s+is perfect here—it uses a positive lookbehind to split after sentence-ending punctuation, ensuring each sentence retains its proper punctuation. - StringBuilder Usage: We use
StringBuilderto build each field, which is far more efficient than repeated string concatenation (avoids unnecessary memory allocations). - Hierarchical Splitting: By handling sentences first, then words, then characters, we maintain logical readability as much as possible before falling back to raw character splitting.
- Edge Case Handling: The code checks if the field length is smaller than the continuation indicator (to avoid invalid outputs) and handles empty lines correctly.
Is There a More Efficient Approach?
Splitting sentences first and processing them individually is indeed the most efficient approach. Any alternative (like scanning the text character-by-character to find split points) would be more complex and likely slower, as regex is optimized for pattern matching, and the layered processing ensures we only scan each part of the text once.
内容的提问来源于stack exchange,提问作者jimbo

