You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Rust中按UTF-8索引对字符串进行切片?

在Rust中按UTF-8字符索引切片字符串

核心问题说明

Rust的&str默认切片基于字节索引,但UTF-8字符可能占用1-4个字节,直接用字节索引切片会破坏编码完整性。你需要的是按字符索引(即第N个UTF-8字符)切片,标准库没有提供直接函数,但可以通过转换字符索引到字节索引实现。

标准库实现方案

单次切片实现

通过char_indices()迭代器获取指定字符索引对应的字节位置,再完成切片:

fn slice_by_char(s: &str, start: usize, end: usize) -> Option<&str> {
    // 边界合法性检查
    if start > end || end > s.chars().count() {
        return None;
    }

    // 获取起始字符的字节偏移
    let start_byte = s.char_indices().nth(start).map(|(pos, _)| pos)?;
    // 获取结束位置的字节偏移(end为排他索引,对应下一个字符的起始位置)
    let end_byte = s.char_indices().nth(end).map(|(pos, _)| pos).unwrap_or(s.len());

    Some(&s[start_byte..end_byte])
}

fn main() {
    let utf8_str = "中文字串";
    // 取第0到第2个字符("中文")
    println!("{:?}", slice_by_char(utf8_str, 0, 2));
    // 取第1到第4个字符("文字串")
    println!("{:?}", slice_by_char(utf8_str, 1, 4));
}

高效缓存方案(适合解析器频繁操作)

如果解析器需要多次按字符索引切片,预先缓存所有字符的字节起始位置可避免重复遍历字符串:

struct Utf8Sliceable {
    source: String,
    char_starts: Vec<usize>,
}

impl Utf8Sliceable {
    fn new(source: String) -> Self {
        // 缓存每个字符的字节起始位置
        let char_starts = source.char_indices().map(|(pos, _)| pos).collect();
        Utf8Sliceable { source, char_starts }
    }

    fn slice(&self, start: usize, end: usize) -> Option<&str> {
        if start > end || end > self.char_starts.len() {
            return None;
        }

        let start_byte = self.char_starts[start];
        let end_byte = if end == self.char_starts.len() {
            self.source.len()
        } else {
            self.char_starts[end]
        };

        Some(&self.source[start_byte..end_byte])
    }
}

fn main() {
    let str_data = Utf8Sliceable::new("中文字串".to_string());
    println!("{:?}", str_data.slice(0, 2)); // "中文"
    println!("{:?}", str_data.slice(2, 4)); // "字串"
}

SWC Input API的处理方式

SWC的StringInput::slice方法基于**字节位置(BytePos)**工作,和直接操作&str的字节切片逻辑一致。若要通过字符索引使用该API,需先将字符索引转换为对应的BytePos:

use swc_common::input::{StringInput, Input};
use swc_common::BytePos;

fn char_index_to_byte_pos(s: &str, index: usize) -> Option<BytePos> {
    s.char_indices().nth(index).map(|(pos, _)| BytePos(pos as u32))
}

fn main() {
    let utf8_str = "中文字串";
    let input = StringInput::new(
        utf8_str,
        BytePos(0),
        BytePos(utf8_str.len() as u32)
    );

    // 按字符索引0到2切片(对应"中文")
    let start_pos = char_index_to_byte_pos(utf8_str, 0).unwrap();
    let end_pos = char_index_to_byte_pos(utf8_str, 2).unwrap();
    println!("{:?}", input.slice(start_pos, end_pos));
}

内容的提问来源于stack exchange,提问作者steven-lie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 05:16:36