You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Rust nom库提取字符串末尾文件名并拆分生成二元组

问题原因及解决办法

核心问题点

  • 字符集遗漏了测试用例中出现的反斜杠\:你当前is_a匹配的字符仅包含大小写字母、数字、.、_、,,但第一个测试用例输入为test\test2,其中的反斜杠不在合法字符范围内,导致解析直接失败,返回默认值("", ""),触发第一个断言报错。
  • many_till逻辑不符合需求:many_till会在第一次匹配到终止规则时就停止,你当前的逻辑会把第一个逗号后的内容识别为后缀,比如输入abc, test.txt, test.txt时,会得到前缀为abc、后缀为test.txt,和预期的前缀abc, test.txt不符。
  • 测试用例笔误:原测试用例中test("bc, def, test.txt")的预期输出写为了("abc, def", "test.txt"),前缀开头多了个a,属于输入笔误,需要修正。

修正后代码

use nom::IResult;
use nom::bytes::complete::take_till;
use nom::character::complete::space0;
use nom::combinator::{all_consuming, opt, recognize};
use nom::multi::many0;
use nom::sequence::{preceded, tuple};
use nom::bytes::complete::tag;

fn test(input: &str) -> (String, String) {
    // 定义合法文件名字符:大小写字母、数字、.、_
    fn filename_char(input: &str) -> IResult<&str, char> {
        nom::character::complete::satisfy(|c| c.is_alphanumeric() || c == '.' || c == '_')(input)
    }
    // 定义完整文件名规则
    fn filename(input: &str) -> IResult<&str, &str> {
        recognize(many0(filename_char))(input)
    }
    // 定义分隔符:前后可带空格的逗号
    fn separator(input: &str) -> IResult<&str, &str> {
        recognize(tuple((space0, tag(","), space0)))(input)
    }

    // 解析逻辑:尝试匹配最后一段 [分隔符 + 文件名],前面所有内容为前缀
    let res: IResult<&str, (Option<&str>, &str)> = all_consuming(
        tuple((
            opt(preceded(
                take_till(|c| c == ','),
                preceded(separator, filename)
            )),
            take_till(|_| false)
        ))
    )(input);

    match res {
        Ok((_, (Some(suffix), prefix))) => (prefix.trim_end_matches(&[',', ' ']).to_string(), suffix.to_string()),
        Ok((_, (None, prefix))) => (prefix.to_string(), "".to_string()),
        Err(_) => (input.to_string(), "".to_string())
    }
}

#[test]
fn test1() {
    assert_eq!(test("test\\test2"), ("test\\test2".to_string(), "".to_string()));
    assert_eq!(test("test2\\a.txt, file1"), ("test2\\a.txt".to_string(), "file1".to_string()));
    assert_eq!(test("abc"), ("abc".to_string(), "".to_string()));
    assert_eq!(test("abc, test.txt"), ("abc".to_string(), "test.txt".to_string()));
    // 修正原测试用例的笔误
    assert_eq!(test("bc, def, test.txt"), ("bc, def".to_string(), "test.txt".to_string()));
    assert_eq!(test("abc, test.txt, test.txt"), ("abc, test.txt".to_string(), "test.txt".to_string()));
}

逻辑说明

优先尝试匹配字符串末尾是否存在「分隔符+合法文件名」的结构,存在就把这部分作为后缀,剩下的内容去掉末尾多余的逗号和空格作为前缀;不存在就把全部输入作为前缀,后缀为空。所有解析失败的异常场景也默认返回全输入作为前缀,避免返回空值。

内容的提问来源于stack exchange,提问作者Just a learner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 04:06:03