You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Rust + nom:如何包装解析器以生成Token?

问题分析与解决

你的代码出现的问题核心是生命周期不匹配和输入类型不一致,导致nom无法正确推导关联类型,以下是具体修正方案:

1. 补全缺失的基础定义

首先补全代码里未定义的new方法,否则无法通过编译:

use nom::{IResult, tag};

#[derive(Debug, PartialEq)]
enum TokenType {
    Foo,
    Bar
}

#[derive(Debug, PartialEq)]
struct Token {
    token_type: TokenType,
    start_offset: usize,
    end_offset: usize,
}

impl Token {
    fn new(token_type: TokenType, start: usize, end: usize) -> Self {
        Token { token_type, start_offset: start, end_offset: end }
    }
}

#[derive(Debug, PartialEq)]
struct LexInput<'a> {
    source: &'a str,
    location: usize,
}

impl<'a> LexInput<'a> {
    fn new(source: &'a str, location: usize) -> Self {
        LexInput { source, location }
    }
}

2. 修正token函数的生命周期与输入逻辑

原代码中parser参数的输入是未绑定生命周期的&str,且与自定义LexInput类型不统一,nom无法处理这种混合输入。我们需要让解析器统一接收LexInput,同时利用nom的trait简化位置计算:

先为LexInput实现nom的必要trait:

use nom::{InputLength, Offset};

impl<'a> InputLength for LexInput<'a> {
    fn input_len(&self) -> usize {
        self.source.input_len()
    }
}

impl<'a> Offset for LexInput<'a> {
    fn offset(&self, second: &Self) -> usize {
        self.source.offset(&second.source)
    }
}

然后重写token函数,统一输入类型并修正生命周期:

fn token<'a>(
    parser: impl Fn(LexInput<'a>) -> IResult<LexInput<'a>, &'a str>,
    token_type: TokenType,
) -> impl Fn(LexInput<'a>) -> IResult<LexInput<'a>, Token> {
    move |input: LexInput<'a>| {
        let start_offset = input.location;
        let (remaining_input, matched) = parser(input)?;
        let end_offset = start_offset + matched.len();
        let token = Token::new(token_type, start_offset, end_offset);
        Ok((remaining_input, token))
    }
}

3. 适配nom内置解析器到LexInput

nom的内置解析器(比如tag)默认处理&str,需要包装成适配LexInput的版本:

fn wrap_tag<'a>(tag_str: &'static str) -> impl Fn(LexInput<'a>) -> IResult<LexInput<'a>, &'a str> {
    move |input: LexInput<'a>| {
        let (remaining_source, matched) = tag(tag_str)(input.source)?;
        let remaining = LexInput::new(remaining_source, input.location + tag_str.len());
        Ok((remaining, matched))
    }
}

4. 测试修正后的代码

现在可以正常调用并验证:

fn main() {
    let (remaining, token) = token(wrap_tag("|"), TokenType::Bar)(LexInput::new("|foo", 0)).unwrap();
    assert_eq!(remaining.source, "foo");
    assert_eq!(remaining.location, 1);
    assert_eq!(token, Token::new(TokenType::Bar, 0, 1));
}

关键问题总结

  • 原代码中parser的输入&str未绑定到'a生命周期,导致nom无法推导一致的关联类型,这是你看到奇怪错误信息的根源。
  • nom要求解析器的输入输出类型保持统一,混合使用&str和自定义LexInput会触发类型推导失败。
  • 为自定义输入类型实现nom的InputLength和Offset trait,可以更便捷地处理位置计算逻辑。

内容的提问来源于stack exchange,提问作者hillin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 04:45:15