能否用CodeQL提取Ruby项目全部正则表达式?求可行方案
提取Ruby项目中所有正则表达式的方案
一、CodeQL 全量提取方案
你当前使用的CodeQL查询是针对ReDoS风险的安全分析查询,只会返回存在指数级回溯风险的正则表达式片段,因此无法获取全量结果。要提取所有正则表达式,需要修改查询逻辑,直接匹配Ruby代码中所有正则表达式节点:
基础全量查询
import codeql.ruby.Regexp from RegExp re select re.getLocation(), "正则表达式内容: " + re.getSource()
该查询会返回项目中所有正则表达式的位置和内容,涵盖两种常见定义方式:
- 字面量形式:
/pattern/、%r{pattern} - 构造函数形式:
Regexp.new("pattern")
进阶:区分正则类型(可选)
如果需要区分不同定义方式的正则,可以细化查询:
import codeql.ruby.Regexp import codeql.ruby.Ast from RegExp re select re.getLocation(), re.getSource(), case re instanceof RegExpLiteral then "字面量定义" re instanceof RegExpConstructorCall then "构造函数定义" else "其他类型" end as 定义方式
二、非CodeQL方案
1. Ruby AST静态分析脚本
使用parser gem解析Ruby代码的抽象语法树(AST),精准定位所有正则表达式节点,支持处理复杂场景(如插值正则/pattern#{var}/)。
步骤:
- 安装依赖:
gem install parser - 运行以下脚本:
require 'parser/current' def extract_regexes(file_path) buffer = Parser::Source::Buffer.new(file_path) buffer.source = File.read(file_path) ast = Parser::CurrentRuby.parse(buffer) regexes = [] processor = Class.new(Parser::AST::Processor) do attr_accessor :regexes def initialize(regexes) @regexes = regexes super() end # 处理字面量正则(/.../、%r{...}) def on_regexp_literal(node) regex_content = node.children.first.source regexes << { type: "字面量", content: regex_content, line: node.loc.line } super(node) end # 处理Regexp.new构造的正则 def on_send(node) if node.method_name == :new && node.receiver&.type == :const && node.receiver.children.first == :Regexp first_arg = node.children[2] if first_arg&.type == :str regexes << { type: "构造函数", content: first_arg.children.first, line: node.loc.line } elsif first_arg&.type == :dstr # 处理动态字符串参数 regexes << { type: "动态构造", content: first_arg.loc.source, line: node.loc.line } end end super(node) end end.new(regexes) processor.process(ast) regexes end # 遍历项目中所有Ruby文件 Dir.glob("./**/*.rb").each do |file| next if File.directory?(file) puts "\n=== #{file} ===" extract_regexes(file).each do |r| puts "行#{r[:line]} [#{r[:type]}]: #{r[:content]}" end end
2. 命令行grep快速扫描
适合不需要精准分析的场景,通过正则匹配常见的Ruby正则写法:
# 匹配字面量正则(包括/.../和%r{...}) grep -rE '/[^/]+/|%r\[[^\]]+\]|%r\{[^\}]+\}' --include="*.rb" ./your-ruby-project # 匹配Regexp.new构造的正则 grep -rE 'Regexp\.new\(["'\''"].*["'\'']\)' --include="*.rb" ./your-ruby-project
⚠️ 局限性:无法区分字符串中的/和真实正则,也无法处理动态生成的正则(如变量拼接的模式)。
内容的提问来源于stack exchange,提问作者Kakashi77
相关产品推荐
相关产品推荐

