如何使用nom解析器统计代码中的注释行数?
使用nom统计代码注释行数的问题解决
问题背景
需要用nom编写解析器,统计文件中两种注释的总行数:
- 单行注释:以
//开头,至行尾结束 - 多行注释:以
/*开头,*/结尾,统计其中包含的行数
现有代码的count_comment_lines逻辑无效,无法正确匹配注释并统计,同时需要替代take_until的解析器驱动工具。
现有代码问题分析
use nom::{ Err, IResult, Parser, branch::alt, bytes::complete::{is_not, tag, take_until}, character::complete::{char, line_ending}, combinator::{value, eof, map, not, success, all_consuming}, error::{ErrorKind, ParseError}, multi::many0, sequence::{pair, tuple, preceded, delimited, terminated}, }; // 匹配单行注释 fn single_line_comment(s: &str) -> IResult<&str,&str> { preceded(tag("//"), is_not("\n\r"))(s) } // 单行注释计数(固定1行) pub fn count_single_line_comment(i: &str) -> IResult<&str, usize> { value(1, single_line_comment)(i) } // 匹配多行注释(存在缺陷) fn multi_line_comments(i: &str) -> IResult<&str, &str> { delimited(tag("/*"), take_until("*/"), tag("*/"))(i) } // 统计多行注释的行数 fn count_multi_line_comments(i: &str) -> IResult<&str, usize> { map(multi_line_comments, |s| s.lines().count())(i) } // 匹配任意一种注释并计数 fn _count_comment_lines(s: &str) -> IResult<&str, usize> { alt((count_single_line_comment, count_multi_line_comments))(s) } // 逻辑错误的统计函数 pub fn count_comment_lines(s: &str) -> IResult<&str, Vec<usize>> { many0( alt((_count_comment_lines, preceded( not(_count_comment_lines), _count_comment_lines ))) )(s) }
核心问题
- 多行注释匹配缺陷:
take_until("*/")会在遇到第一个*/片段时提前终止,若注释内包含类似*/*/的内容会导致匹配错误,且它是基于字节的匹配,而非解析器驱动。 - 统计逻辑错误:
count_comment_lines中的preceded(not(_count_comment_lines), _count_comment_lines)不会消耗非注释内容(not仅做断言不消耗输入),会进入死循环;同时缺少对非注释内容的跳过逻辑,无法正确遍历整个文本。
解决方案
1. 修复多行注释解析
替换take_until为解析器驱动的组合,确保只有完整的*/才会终止匹配:
use nom::multi::many0; use nom::branch::alt; use nom::bytes::complete::is_not; use nom::combinator::recognize; use nom::sequence::pair; // 正确匹配多行注释,避免提前终止 fn multi_line_comments(i: &str) -> IResult<&str, &str> { recognize( pair( tag("/*"), pair( many0(alt(( is_not("*"), // 匹配非*的任意字符 tag("*").and(not(tag("/"))) // 匹配单独的*,排除*/ ))), tag("*/") ) ) )(i) }
2. 实现非注释内容跳过逻辑
编写解析器跳过所有非注释开头的内容:
// 跳过所有非注释的内容(直到遇到//、/*或EOF) fn skip_non_comment(i: &str) -> IResult<&str, ()> { let (i, _) = many0(alt(( is_not("/"), // 匹配非/的任意字符 tag("/").and(not(alt((tag("/"), tag("*"))))) // 匹配单独的/,排除//和/* )))(i)?; Ok((i, ())) }
3. 重构统计函数
重新设计逻辑:跳过非注释内容 → 匹配注释并计数 → 重复此过程直到文本结束,最后求和所有计数:
// 统计所有注释的总行数 pub fn count_comment_lines(s: &str) -> IResult<&str, usize> { let (s, counts) = many0( preceded( skip_non_comment, alt((count_single_line_comment, count_multi_line_comments)) ) )(s)?; // 跳过剩余的非注释内容 let (s, _) = skip_non_comment(s)?; Ok((s, counts.into_iter().sum())) }
如果需要保留每个注释的行数列表,只需修改返回值为Vec<usize>:
pub fn collect_comment_lines(s: &str) -> IResult<&str, Vec<usize>> { let (s, counts) = many0( preceded( skip_non_comment, alt((count_single_line_comment, count_multi_line_comments)) )(s)?; let (s, _) = skip_non_comment(s)?; Ok((s, counts)) }
测试验证
针对测试样本的预期结果:
SAMPLE1:2行(两个单行注释)SAMPLE2:3行(多行注释包含3行文本)SAMPLE3:4行(1行单行注释 + 3行多行注释)SAMPLE4:3行(三个单行注释,其中//Second//NotThird算1行)
使用all_consuming包裹统计函数可以确保整个文本被解析:
#[test] fn test_sample1() { assert_eq!(all_consuming(count_comment_lines)(SAMPLE1), Ok(("", 2))); } #[test] fn test_sample2() { assert_eq!(all_consuming(count_comment_lines)(SAMPLE2), Ok(("", 3))); } #[test] fn test_sample3() { assert_eq!(all_consuming(count_comment_lines)(SAMPLE3), Ok(("", 4))); } #[test] fn test_sample4() { assert_eq!(all_consuming(count_comment_lines)(SAMPLE4), Ok(("", 3))); }
内容的提问来源于stack exchange,提问作者red-swan
相关产品推荐
相关产品推荐

