如何将Swift数组分块算法泛化至任意Collection/Sequence(含String)?
Awesome question! Let's walk through how to generalize that chunking logic to work with any Collection or Sequence—including String, which is a common use case folks run into.
First, let's recap why the original Array-only implementation doesn't translate directly: it relies on integer indices and random access (via stride and range subscripting with Int values), which aren't universal across all collections. For example, String uses String.Index (not integers) to reference characters, and not all sequences/collections support random jumps through elements.
Step 1: Generalize for Collection
All collections have stable indices and support subscripting with ranges, so we can build a flexible implementation that works for every type conforming to Collection (including String, Array, Set, and custom collections).
Universal Collection Implementation
This version works for any collection, even those without random access (like linked lists or String):
extension Collection { func chunked(into size: Int) -> [SubSequence] { precondition(size > 0, "Chunk size must be greater than 0") var chunks: [SubSequence] = [] var currentStart = startIndex while currentStart < endIndex { // Calculate the end index for the current chunk (don't go past the collection's end) let currentEnd = index(currentStart, offsetBy: size, limitedBy: endIndex) ?? endIndex chunks.append(self[currentStart..<currentEnd]) currentStart = currentEnd } return chunks } }
Optimized Version for RandomAccessCollection
For collections that support O(1) index operations (like Array, ContiguousArray), we can optimize using stride for better performance:
extension RandomAccessCollection { func chunked(into size: Int) -> [SubSequence] { precondition(size > 0, "Chunk size must be greater than 0") return stride(from: startIndex, to: endIndex, by: size).map { start in let end = index(start, offsetBy: size, limitedBy: endIndex) ?? endIndex return self[start..<end] } } }
Step 2: Generalize for Sequence
Sequences don't have stable indices and can only be traversed once, so we need a different approach—we'll collect elements into temporary arrays until we hit the chunk size:
extension Sequence { func chunked(into size: Int) -> [[Element]] { precondition(size > 0, "Chunk size must be greater than 0") var chunks: [[Element]] = [] var currentChunk: [Element] = [] for element in self { currentChunk.append(element) if currentChunk.count == size { chunks.append(currentChunk) currentChunk = [] } } // Add any remaining elements as the final chunk if !currentChunk.isEmpty { chunks.append(currentChunk) } return chunks } }
Using with String
Since String conforms to Collection (its elements are Character values), you can use the Collection extension directly to chunk by characters:
let testString = "Hello, let's chunk this string into 5-character pieces!" let charChunks = testString.chunked(into: 5) // Result: ["Hello", ", let", "'s ch", "unk t", "his s", "tring", " into", " 5-ch", "aract", "er pi", "eces!"]
If you need to chunk by UTF-8 bytes or Unicode scalars instead, just access the corresponding view and call chunked(into:):
let utf8Chunks = testString.utf8.chunked(into: 5) let scalarChunks = testString.unicodeScalars.chunked(into: 5)
Key Notes
- The
Collectionextensions return[SubSequence]instead of[[Element]]—this avoids copying elements unnecessarily (e.g.,SubstringforStringis a lightweight view into the original string). - We add a
preconditionto guard against invalid chunk sizes (0 or negative), which makes the code more robust. - The
Sequenceimplementation returns[[Element]]because sequences don't support subviews—we have to copy elements into new arrays.
内容的提问来源于stack exchange,提问作者blackjacx

