如何用Superpower Parser解析含末尾固定关键词的可变前缀字符串
(?'prefix'.*)(some\skeyword)的解析器 需求说明
需要实现一个解析器,功能等价于正则表达式 (?'prefix'.*)(some\skeyword):从输入字符串中提取最后出现的"some keyword"之前的所有内容作为前缀。例如输入字符串"This is some prefix of some keyword",需提取出前缀"This is some prefix of"。
问题分析
你当前的代码中,Text.Many()是贪婪匹配,会消耗所有输入Token(包括组成Keyword的"some"和"keyword"),导致后续Keyword匹配时输入已耗尽,抛出expected "some" but end of input is encountered错误。
解决方案
核心思路是使用**否定前瞻(Negative Lookahead)**限制前缀文本的匹配范围,确保当前匹配的文本Token之后不会立即出现目标关键词序列。
修改后的完整代码
using Superpower; using Superpower.Model; // 假设StoryToken是你定义的Token类型枚举 public enum StoryToken { Text } public class StoryParsers { // 匹配单个文本Token public static readonly TokenListParser<StoryToken, string> Text = Token.EqualTo(StoryToken.Text) .Select(x => x.Span.ToStringValue()); // 匹配不区分大小写的"some keyword"序列 public static readonly TokenListParser<StoryToken, string> Keyword = from someToken in Token.EqualToValueIgnoreCase(StoryToken.Text, "some") from keywordToken in Token.EqualToValueIgnoreCase(StoryToken.Text, "keyword") select "some keyword"; // 匹配前缀文本:单个Text Token,且后续不会出现Keyword序列 private static readonly TokenListParser<StoryToken, string> PrefixText = Text.Before(Keyword.Not()); // 最终解析器:提取前缀和关键词 public static readonly TokenListParser<StoryToken, (string Prefix, string Keyword)> TextWithKeyword = from prefixParts in PrefixText.Many() from keyword in Keyword select ( Prefix: string.Join(" ", prefixParts), Keyword: keyword ); }
代码说明
PrefixText解析器:
使用Before(Keyword.Not())实现否定前瞻,确保当前匹配的TextToken之后,不会立即出现完整的Keyword序列。这样PrefixText.Many()只会消耗到Keyword之前的所有Token,不会吃掉关键词的起始Token。最终解析器返回值:
返回元组(string Prefix, string Keyword),相比直接拼接字符串更便于后续处理;如果需要拼接成原格式,可修改为select $"{string.Join(" ", prefixParts)} {keyword}"。
验证效果
对输入字符串"This is some prefix of some keyword"分词后,解析器会正确提取:
- 前缀:
"This is some prefix of" - 关键词:
"some keyword"
内容的提问来源于stack exchange,提问作者Kirill Gribunin

