如何在Rust中按UTF-8索引对字符串进行切片?
在Rust中按UTF-8字符索引切片字符串
核心问题说明
Rust的&str默认切片基于字节索引,但UTF-8字符可能占用1-4个字节,直接用字节索引切片会破坏编码完整性。你需要的是按字符索引(即第N个UTF-8字符)切片,标准库没有提供直接函数,但可以通过转换字符索引到字节索引实现。
标准库实现方案
单次切片实现
通过char_indices()迭代器获取指定字符索引对应的字节位置,再完成切片:
fn slice_by_char(s: &str, start: usize, end: usize) -> Option<&str> { // 边界合法性检查 if start > end || end > s.chars().count() { return None; } // 获取起始字符的字节偏移 let start_byte = s.char_indices().nth(start).map(|(pos, _)| pos)?; // 获取结束位置的字节偏移(end为排他索引,对应下一个字符的起始位置) let end_byte = s.char_indices().nth(end).map(|(pos, _)| pos).unwrap_or(s.len()); Some(&s[start_byte..end_byte]) } fn main() { let utf8_str = "中文字串"; // 取第0到第2个字符("中文") println!("{:?}", slice_by_char(utf8_str, 0, 2)); // 取第1到第4个字符("文字串") println!("{:?}", slice_by_char(utf8_str, 1, 4)); }
高效缓存方案(适合解析器频繁操作)
如果解析器需要多次按字符索引切片,预先缓存所有字符的字节起始位置可避免重复遍历字符串:
struct Utf8Sliceable { source: String, char_starts: Vec<usize>, } impl Utf8Sliceable { fn new(source: String) -> Self { // 缓存每个字符的字节起始位置 let char_starts = source.char_indices().map(|(pos, _)| pos).collect(); Utf8Sliceable { source, char_starts } } fn slice(&self, start: usize, end: usize) -> Option<&str> { if start > end || end > self.char_starts.len() { return None; } let start_byte = self.char_starts[start]; let end_byte = if end == self.char_starts.len() { self.source.len() } else { self.char_starts[end] }; Some(&self.source[start_byte..end_byte]) } } fn main() { let str_data = Utf8Sliceable::new("中文字串".to_string()); println!("{:?}", str_data.slice(0, 2)); // "中文" println!("{:?}", str_data.slice(2, 4)); // "字串" }
SWC Input API的处理方式
SWC的StringInput::slice方法基于**字节位置(BytePos)**工作,和直接操作&str的字节切片逻辑一致。若要通过字符索引使用该API,需先将字符索引转换为对应的BytePos:
use swc_common::input::{StringInput, Input}; use swc_common::BytePos; fn char_index_to_byte_pos(s: &str, index: usize) -> Option<BytePos> { s.char_indices().nth(index).map(|(pos, _)| BytePos(pos as u32)) } fn main() { let utf8_str = "中文字串"; let input = StringInput::new( utf8_str, BytePos(0), BytePos(utf8_str.len() as u32) ); // 按字符索引0到2切片(对应"中文") let start_pos = char_index_to_byte_pos(utf8_str, 0).unwrap(); let end_pos = char_index_to_byte_pos(utf8_str, 2).unwrap(); println!("{:?}", input.slice(start_pos, end_pos)); }
内容的提问来源于stack exchange,提问作者steven-lie
相关产品推荐
相关产品推荐

