Roslyn代码生成器中高效解析特定脚本文本的最优方案
我最终采用了正则表达式方案,它速度很快,因此我一直沿用该方案。
需求说明
输入文本
@class using System.Text.Json @attribute using System @attribute using System.Text @attribute required Type theType @attribute optional bool IsReadOnly = false @attribute optional bool SomethingElse
目标生成对象
new UsingClause(Target.Class, "System.Text.Json"); new UsingClause(Target.Attribute, "System"); new UsingClause(Target.Attribute, "System.Text"); new AttributeProperty(required: true, name: "theType", type: "Type", default: null); new AttributeProperty(required: false, name: "IsReadOnly", type: "bool", default: "true"); new AttributeProperty(required: false, name: "SomethingElse", type: "bool", default: null);
核心问题
我原本习惯用字符串分割成数组的方式处理这类文本,但因为这段逻辑要用于Roslyn代码生成器,要求极致高效。目前我尝试用StringReader逐行读取文本,配合预编译的正则表达式做匹配,但对这种场景下的高效处理标准不太熟悉,担心当前的实现方式有问题。
当前使用的正则代码
private readonly static Regex Regex = new Regex( pattern: @"^\s*((@attribute)\s+(using)\s+(.*))|((@attribute)\s+(optional|required)\s+(\w+[\w\.]*)\s+(\w+)(\s*\=\s*(.*))?)|((@class)\s+(using)\s+(.*))\s*$", options: RegexOptions.IgnoreCase | RegexOptions.Multiline | RegexOptions.Compiled);
方案分析与优化建议
正则方案的合理性
在Roslyn代码生成的场景下,预编译正则(带RegexOptions.Compiled)是完全符合高效要求的选择——编译后的正则会生成原生IL代码,执行效率远高于解释型正则。搭配StringReader逐行处理,既能避免一次性加载大文本到内存,也能保证每一行的处理速度。
正则写法优化
当前的正则虽然能完成匹配,但存在冗余捕获组、匹配逻辑复杂的问题,可从以下几点优化,进一步提升效率和可读性:
- 拆分正则逻辑:把三种不同类型的行(类引用、属性引用、属性定义)拆成三个独立的预编译正则,逐行匹配时先判断行类型再对应处理,每个正则的逻辑更简单,匹配速度更快。
- 精简捕获组:用非捕获组
(?:...)替代不需要提取数据的分组,只保留需要获取的内容对应的捕获组,减少匹配时的内存开销。 - 明确匹配范围:比如把
\w+[\w\.]*简化为[\w.]+;默认值部分用([^#\r\n]*)替代(.*),避免意外跨行匹配(如果输入有注释的话)。
优化后的正则示例:
// 匹配@class using 行 private static readonly Regex ClassUsingRegex = new Regex( @"^\s*@class\s+using\s+([\w.]+)\s*$", RegexOptions.IgnoreCase | RegexOptions.Compiled); // 匹配@attribute using 行 private static readonly Regex AttrUsingRegex = new Regex( @"^\s*@attribute\s+using\s+([\w.]+)\s*$", RegexOptions.IgnoreCase | RegexOptions.Compiled); // 匹配@attribute 属性定义行 private static readonly Regex AttrPropertyRegex = new Regex( @"^\s*@attribute\s+(required|optional)\s+([\w.]+)\s+(\w+)(?:\s*\=\s*([^#\r\n]*))?\s*$", RegexOptions.IgnoreCase | RegexOptions.Compiled);
其他高效处理细节
- 逐行读取时直接跳过空行或注释行(比如以
#开头的行),避免无效匹配。 - 处理字符串时优先用
Span<char>替代string的分割、截取操作,减少内存分配。 - 如果需要创建大量目标对象,可以提前初始化对象池,减少GC压力。
内容的提问来源于stack exchange,提问作者Peter Morris
相关产品推荐
相关产品推荐

