You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用nom解析器统计代码中的注释行数?

使用nom统计代码注释行数的问题解决

问题背景

需要用nom编写解析器,统计文件中两种注释的总行数:

  • 单行注释:以//开头,至行尾结束
  • 多行注释:以/*开头,*/结尾,统计其中包含的行数

现有代码的count_comment_lines逻辑无效,无法正确匹配注释并统计,同时需要替代take_until的解析器驱动工具。

现有代码问题分析

use nom::{
  Err, IResult, Parser,
  branch::alt,
  bytes::complete::{is_not, tag, take_until},
  character::complete::{char, line_ending},
  combinator::{value, eof, map, not, success, all_consuming},
  error::{ErrorKind, ParseError},
  multi::many0,
  sequence::{pair, tuple, preceded, delimited, terminated},
};

// 匹配单行注释
fn single_line_comment(s: &str) -> IResult<&str,&str> {
  preceded(tag("//"), is_not("\n\r"))(s)
}

// 单行注释计数(固定1行)
pub fn count_single_line_comment(i: &str) -> IResult<&str, usize> {
  value(1, single_line_comment)(i)
}

// 匹配多行注释(存在缺陷)
fn multi_line_comments(i: &str) -> IResult<&str, &str> {
  delimited(tag("/*"), take_until("*/"), tag("*/"))(i)
}

// 统计多行注释的行数
fn count_multi_line_comments(i: &str) -> IResult<&str, usize> {
  map(multi_line_comments, |s| s.lines().count())(i) 
}

// 匹配任意一种注释并计数
fn _count_comment_lines(s: &str) -> IResult<&str, usize> {
  alt((count_single_line_comment, count_multi_line_comments))(s)
}

// 逻辑错误的统计函数
pub fn count_comment_lines(s: &str) -> IResult<&str, Vec<usize>> {
    many0(
      alt((_count_comment_lines, 
           preceded(
            not(_count_comment_lines),
           _count_comment_lines
          )))
    )(s)
}

核心问题

  1. 多行注释匹配缺陷:take_until("*/")会在遇到第一个*/片段时提前终止,若注释内包含类似*/*/的内容会导致匹配错误,且它是基于字节的匹配,而非解析器驱动。
  2. 统计逻辑错误:count_comment_lines中的preceded(not(_count_comment_lines), _count_comment_lines)不会消耗非注释内容(not仅做断言不消耗输入),会进入死循环;同时缺少对非注释内容的跳过逻辑,无法正确遍历整个文本。

解决方案

1. 修复多行注释解析

替换take_until为解析器驱动的组合,确保只有完整的*/才会终止匹配:

use nom::multi::many0;
use nom::branch::alt;
use nom::bytes::complete::is_not;
use nom::combinator::recognize;
use nom::sequence::pair;

// 正确匹配多行注释,避免提前终止
fn multi_line_comments(i: &str) -> IResult<&str, &str> {
    recognize(
        pair(
            tag("/*"),
            pair(
                many0(alt((
                    is_not("*"), // 匹配非*的任意字符
                    tag("*").and(not(tag("/"))) // 匹配单独的*,排除*/
                ))),
                tag("*/")
            )
        )
    )(i)
}

2. 实现非注释内容跳过逻辑

编写解析器跳过所有非注释开头的内容:

// 跳过所有非注释的内容(直到遇到//、/*或EOF)
fn skip_non_comment(i: &str) -> IResult<&str, ()> {
    let (i, _) = many0(alt((
        is_not("/"), // 匹配非/的任意字符
        tag("/").and(not(alt((tag("/"), tag("*"))))) // 匹配单独的/,排除//和/*
    )))(i)?;
    Ok((i, ()))
}

3. 重构统计函数

重新设计逻辑:跳过非注释内容 → 匹配注释并计数 → 重复此过程直到文本结束,最后求和所有计数:

// 统计所有注释的总行数
pub fn count_comment_lines(s: &str) -> IResult<&str, usize> {
    let (s, counts) = many0(
        preceded(
            skip_non_comment,
            alt((count_single_line_comment, count_multi_line_comments))
        )
    )(s)?;
    // 跳过剩余的非注释内容
    let (s, _) = skip_non_comment(s)?;
    Ok((s, counts.into_iter().sum()))
}

如果需要保留每个注释的行数列表,只需修改返回值为Vec<usize>:

pub fn collect_comment_lines(s: &str) -> IResult<&str, Vec<usize>> {
    let (s, counts) = many0(
        preceded(
            skip_non_comment,
            alt((count_single_line_comment, count_multi_line_comments))
        )(s)?;
    let (s, _) = skip_non_comment(s)?;
    Ok((s, counts))
}

测试验证

针对测试样本的预期结果:

  • SAMPLE1:2行(两个单行注释)
  • SAMPLE2:3行(多行注释包含3行文本)
  • SAMPLE3:4行(1行单行注释 + 3行多行注释)
  • SAMPLE4:3行(三个单行注释,其中//Second//NotThird算1行)

使用all_consuming包裹统计函数可以确保整个文本被解析:

#[test]
fn test_sample1() {
    assert_eq!(all_consuming(count_comment_lines)(SAMPLE1), Ok(("", 2)));
}

#[test]
fn test_sample2() {
    assert_eq!(all_consuming(count_comment_lines)(SAMPLE2), Ok(("", 3)));
}

#[test]
fn test_sample3() {
    assert_eq!(all_consuming(count_comment_lines)(SAMPLE3), Ok(("", 4)));
}

#[test]
fn test_sample4() {
    assert_eq!(all_consuming(count_comment_lines)(SAMPLE4), Ok(("", 3)));
}

内容的提问来源于stack exchange,提问作者red-swan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 04:07:31