Go语言中rune切片长度与utf8.RuneCountInString的区别是什么?
Hey there! Great question—these two approaches might seem similar at first glance, but they have important differences in memory usage, performance, and intended use cases. Let’s break them down:
1. Memory Overhead
When you convert a string to a []rune slice like this:
s := "世界" runes := []rune(s) fmt.Println(len(runes)) // Outputs 2
Go allocates a new slice in memory that stores every decoded rune from the string. For short strings this is trivial, but for large UTF-8 strings (think multi-kilobyte text), this creates an extra copy of all the rune data, which can eat up unnecessary memory.
On the other hand, utf8.RuneCountInString(s) doesn’t allocate any additional memory for storing runes. It just iterates through the string’s bytes, counts valid UTF-8 sequences, and returns the total—no extra slice created.
2. Performance
Because of the memory allocation, converting to a []rune and checking its length is generally slower than using utf8.RuneCountInString, especially for large strings. The standard library’s RuneCountInString is optimized to count runes efficiently without the overhead of copying data into a slice.
If you benchmark both methods, you’ll see a noticeable difference: RuneCountInString will outperform the []rune approach by a significant margin when you only need the count.
3. Use Case Fit
- Use
[]rune(s)when you actually need to work with the individual runes (e.g., modifying characters, iterating over them with index access, or passing them to functions that expect a rune slice). The length check is just a side effect here—you’re getting the slice for other purposes. - Use
utf8.RuneCountInString(s)when your sole goal is to count the number of Unicode characters in the string. It’s the most efficient and idiomatic way to do this in Go if you don’t need the rune data itself.
4. Under-the-Hood Behavior
Both methods decode UTF-8 sequences to count runes, but:
[]rune(s)decodes every rune and stores them in a contiguous slice. This means you’re doing extra work to store data you might not need if you only care about the count.utf8.RuneCountInStringiterates through the string’s bytes, usesutf8.DecodeRuneInStringto skip over each valid UTF-8 sequence, and increments a counter each time. No storage of runes—just pure counting.
Example Benchmark (for context)
Here’s a quick benchmark to illustrate the performance gap:
package main import ( "testing" "unicode/utf8" ) var testString = "こんにちは世界! Hello World! 🌍" func BenchmarkRuneSliceLen(b *testing.B) { for i := 0; i < b.N; i++ { _ = len([]rune(testString)) } } func BenchmarkRuneCountInString(b *testing.B) { for i := 0; i < b.N; i++ { _ = utf8.RuneCountInString(testString) } }
Running this would show BenchmarkRuneCountInString runs far more operations per second than BenchmarkRuneSliceLen, all thanks to avoiding unnecessary memory allocation.
内容的提问来源于stack exchange,提问作者Wang Ke

