You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Rust遍历字符解析字符串时索引偏移,请求排查逻辑缺陷

1brc挑战Rust解析问题:Reykjavík城市名截断导致温度解析失败

我正在通过1brc挑战学习Rust,采用分块并行读取文件的方案——将文件块以Box<[u8]>通过通道传递,转换为字符串后遍历字符构建城市与温度的映射表。但解析代码在处理Reykjavík时持续失败:end_idx偏移一位,导致城市名变为"Reykjaví",后续浮点数解析也跟着出错。

相关代码如下:

const VALUE_SEPARATOR: char = ';';
const NEW_LINE: char = '\n'; 

//constants declared at the top   

for chunk in chunk_receiver {
        let chunk_string = std::str::from_utf8(&chunk).unwrap();

        let mut start_idx = 0;
        let mut end_idx = 0;
    
        for (idx, elem) in chunk_string.chars().enumerate() {
            match elem {
                VALUE_SEPARATOR => {
                    end_idx = idx;
                },
                NEW_LINE => {
                    let city = chunk_string[start_idx..end_idx].to_string();
                    start_idx = end_idx + 1;
                    

                    if (idx - end_idx) > 1 && city.len() > 0 {
                        let temperature: f32 = fast_float::parse::<f32, _>(chunk_string[start_idx..idx].as_bytes()).unwrap();
                        start_idx = idx + 1;

                        map.entry(city)
                            .and_modify(|metric| metric.update(temperature))
                            .or_insert_with(|| Metrics::new(temperature));
                    }
                },
                _ => {
                    continue;
                },
            };
        }
    }

问题根源

核心问题是字符索引与字节索引不匹配:Reykjavík里的í是UTF-8多字节字符(占2字节),chunk_string.chars().enumerate()返回的idx是字符的逻辑位置,但字符串切片chunk_string[start_idx..end_idx]用的是字节偏移量。当遇到多字节字符时,两者错位,导致end_idx指向í的第二个字节,直接截断了后面的k。

修复方案

直接处理字节流(更符合1brc的性能要求),避免字符索引和字节索引的混淆:

const VALUE_SEPARATOR: u8 = b';';
const NEW_LINE: u8 = b'\n';

// 保留原有其他逻辑,修改块解析部分
for chunk in chunk_receiver {
    let mut start_idx = 0;
    let mut end_idx = 0;

    // 直接遍历字节数组的索引和字节值
    for (idx, &byte) in chunk.iter().enumerate() {
        match byte {
            VALUE_SEPARATOR => {
                end_idx = idx;
            }
            NEW_LINE => {
                // 从字节切片直接转字符串,确保UTF-8合法性
                let city = std::str::from_utf8(&chunk[start_idx..end_idx]).unwrap().to_string();
                start_idx = end_idx + 1;

                if (idx - end_idx) > 1 && !city.is_empty() {
                    // 直接传递字节切片给fast_float,无需转字符串
                    let temperature: f32 = fast_float::parse::<f32, _>(&chunk[start_idx..idx]).unwrap();
                    start_idx = idx + 1;

                    map.entry(city)
                        .and_modify(|metric| metric.update(temperature))
                        .or_insert_with(|| Metrics::new(temperature));
                }
            }
            _ => continue,
        }
    }
}

额外优化建议

  1. 分块时避免截断多字节字符:如果分块是手动切割的,要确保块的结束位置是换行符,避免把一个UTF-8字符拆到两个块里。
  2. 减少字符串转换:尽量直接操作字节数组,只在需要城市名字符串时再转换,降低内存开销。

内容的提问来源于stack exchange,提问作者Mechanik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 15:32:37