C#分割文本后获取各分段起始位置的高效实现方案咨询
Great question! Your current approach works, but it can get inefficient with large text since it’s doing redundant lookups of entire segment strings. Let’s go through faster, more optimized solutions that avoid extra work:
1. One-Pass Separator Location Scan (Best for Most Cases)
Instead of splitting first and then searching for each segment, scan the original text once to find all occurrences of your separator, then calculate segment start positions directly. This cuts down on repeated string matching and avoids the overhead of splitting first.
Here’s how to implement it:
string separator = Environment.NewLine + "##"; int separatorLength = separator.Length; List<int> segmentStarts = new List<int> { 0 }; // First segment starts at index 0 int currentPos = 0; while ((currentPos = text.IndexOf(separator, currentPos)) != -1) { // Next segment starts right after the end of the separator segmentStarts.Add(currentPos + separatorLength); currentPos += separatorLength; // Skip past the separator to avoid re-matching } // Convert to array if needed int[] pos = segmentStarts.ToArray();
Why this is better:
- Only traverses the text once to find separators, instead of multiple passes for each segment.
- Avoids matching entire segment strings (which can be long) — it only looks for the fixed-length separator, which is faster.
- Works seamlessly with your original split logic (the segments will align perfectly with the start positions).
2. Span Optimization (For Large/Very Large Text)
If you’re working with extremely large text (think megabytes or more), using Span<char> will eliminate unnecessary string allocations and reduce GC pressure, making the scan even faster.
ReadOnlySpan<char> textSpan = text.AsSpan(); ReadOnlySpan<char> separatorSpan = (Environment.NewLine + "##").AsSpan(); List<int> segmentStarts = new List<int> { 0 }; int currentPos = 0; while ((currentPos = textSpan.IndexOf(separatorSpan, currentPos)) != -1) { segmentStarts.Add(currentPos + separatorSpan.Length); currentPos += separatorSpan.Length; } int[] pos = segmentStarts.ToArray();
Why this is better:
Span<char>operates directly on the text’s underlying memory without copying data.- No extra string objects are created during the scan, which is critical for avoiding GC pauses with large datasets.
How Your Original Approach Compares
Your current method (text.IndexOf(slides[i], pos[i-1] + 1)) has two main drawbacks:
- It requires searching for the entire segment string each time, which is slower than searching for a short, fixed separator.
- It effectively does two passes over the text: one for splitting, then another for locating each segment. The one-pass scan above combines these steps into a single efficient pass.
内容的提问来源于stack exchange,提问作者Ahmad

