如何用nom解析指定长度输入?以日期时间串解析为例
用Nom解析固定长度数字段的正确方法
问题背景
需要解析类似(D:20230416140523Z00'00)的字符串,从中提取YYYYMMDDHHmmSS格式的日期时间字段,每个字段是固定长度的数字(比如YYYY是4位、MM是2位等)。尝试用take(n)结合u16解析固定长度前缀时遇到编译错误。
错误原因分析
原代码中使用flat_map(take(n), u16::<_, ()>)的核心问题:
flat_map的第二个参数要求是接收第一个解析结果、返回新Parser的函数,但直接传入u16::<_, ()>会立即执行该解析器,返回的是Result类型,而非符合要求的Parser闭包。- 我们实际需求是将
take(n)获取的字节片段转换为数字类型,而非用新Parser继续解析剩余输入,因此应该使用map_res而非flat_map。
正确实现方式
方法1:基础版固定长度数字解析
使用map_res将固定长度字节片段转换为数字,它支持处理转换过程中的错误(比如非数字字符、长度不匹配):
use nom::bytes::complete::take; use nom::combinator::map_res; use nom::error::Error; type IResult<'a, T> = nom::IResult<&'a [u8], T, Error<&'a [u8]>>; fn parse_fixed_length_prefix<'a>(input: &[u8], n: usize) -> IResult<u16> { map_res(take(n), |bytes: &[u8]| { // 先转UTF-8字符串,再解析为u16 std::str::from_utf8(bytes).and_then(|s| s.parse::<u16>()) })(input) } #[test] fn test() { let output = parse_fixed_length_prefix(b"123456789", 4); assert_eq!(output, Ok((b"56789", 1234u16))); }
方法2:封装通用解析器(适合多字段场景)
如果需要多次解析不同长度的数字段,可以封装一个通用的固定长度数字解析器,支持u16/u8/u32等多种数字类型:
use nom::bytes::complete::take; use nom::combinator::map_res; use nom::error::Error; type IResult<'a, T> = nom::IResult<&'a [u8], T, Error<&'a [u8]>>; // 通用固定长度数字解析器,支持任意可从字符串解析的数字类型 fn fixed_num<'a, N: std::str::FromStr>(n: usize) -> impl Fn(&'a [u8]) -> IResult<N> where <N as std::str::FromStr>::Err: std::fmt::Display, { move |input| { map_res(take(n), |bytes: &[u8]| { std::str::from_utf8(bytes).and_then(|s| s.parse::<N>()) })(input) } } // 解析完整YYYYMMDDHHmmSS格式的示例 fn parse_datetime(input: &[u8]) -> IResult<(u16, u8, u8, u8, u8, u8)> { let (input, year) = fixed_num::<u16>(4)(input)?; let (input, month) = fixed_num::<u8>(2)(input)?; let (input, day) = fixed_num::<u8>(2)(input)?; let (input, hour) = fixed_num::<u8>(2)(input)?; let (input, minute) = fixed_num::<u8>(2)(input)?; let (input, second) = fixed_num::<u8>(2)(input)?; Ok((input, (year, month, day, hour, minute, second))) } #[test] fn test_datetime() { let output = parse_datetime(b"20230416140523Z00'00"); assert_eq!(output, Ok((b"Z00'00", (2023, 4, 16, 14, 5, 23)))); }
补充说明
- 用
Error<&[u8]>替代空错误类型(),能提供更详细的错误信息,便于调试。 - 日期时间数字串通常是合法UTF-8,若需处理非UTF-8场景,可根据实际需求调整字节转数字的逻辑。
内容的提问来源于stack exchange,提问作者George
相关产品推荐
相关产品推荐

