You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否用CodeQL提取Ruby项目全部正则表达式?求可行方案

提取Ruby项目中所有正则表达式的方案

一、CodeQL 全量提取方案

你当前使用的CodeQL查询是针对ReDoS风险的安全分析查询,只会返回存在指数级回溯风险的正则表达式片段,因此无法获取全量结果。要提取所有正则表达式,需要修改查询逻辑,直接匹配Ruby代码中所有正则表达式节点:

基础全量查询

import codeql.ruby.Regexp

from RegExp re
select re.getLocation(), "正则表达式内容: " + re.getSource()

该查询会返回项目中所有正则表达式的位置和内容,涵盖两种常见定义方式:

  • 字面量形式:/pattern/、%r{pattern}
  • 构造函数形式:Regexp.new("pattern")

进阶:区分正则类型(可选)

如果需要区分不同定义方式的正则,可以细化查询:

import codeql.ruby.Regexp
import codeql.ruby.Ast

from RegExp re
select 
  re.getLocation(),
  re.getSource(),
  case
    re instanceof RegExpLiteral then "字面量定义"
    re instanceof RegExpConstructorCall then "构造函数定义"
    else "其他类型"
  end as 定义方式

二、非CodeQL方案

1. Ruby AST静态分析脚本

使用parser gem解析Ruby代码的抽象语法树(AST),精准定位所有正则表达式节点,支持处理复杂场景(如插值正则/pattern#{var}/)。

步骤:

  1. 安装依赖:gem install parser
  2. 运行以下脚本:
require 'parser/current'

def extract_regexes(file_path)
  buffer = Parser::Source::Buffer.new(file_path)
  buffer.source = File.read(file_path)
  ast = Parser::CurrentRuby.parse(buffer)

  regexes = []

  processor = Class.new(Parser::AST::Processor) do
    attr_accessor :regexes

    def initialize(regexes)
      @regexes = regexes
      super()
    end

    # 处理字面量正则(/.../、%r{...})
    def on_regexp_literal(node)
      regex_content = node.children.first.source
      regexes << { type: "字面量", content: regex_content, line: node.loc.line }
      super(node)
    end

    # 处理Regexp.new构造的正则
    def on_send(node)
      if node.method_name == :new && node.receiver&.type == :const && node.receiver.children.first == :Regexp
        first_arg = node.children[2]
        if first_arg&.type == :str
          regexes << { type: "构造函数", content: first_arg.children.first, line: node.loc.line }
        elsif first_arg&.type == :dstr # 处理动态字符串参数
          regexes << { type: "动态构造", content: first_arg.loc.source, line: node.loc.line }
        end
      end
      super(node)
    end
  end.new(regexes)

  processor.process(ast)
  regexes
end

# 遍历项目中所有Ruby文件
Dir.glob("./**/*.rb").each do |file|
  next if File.directory?(file)
  puts "\n=== #{file} ==="
  extract_regexes(file).each do |r|
    puts "行#{r[:line]} [#{r[:type]}]: #{r[:content]}"
  end
end

2. 命令行grep快速扫描

适合不需要精准分析的场景,通过正则匹配常见的Ruby正则写法:

# 匹配字面量正则(包括/.../和%r{...})
grep -rE '/[^/]+/|%r\[[^\]]+\]|%r\{[^\}]+\}' --include="*.rb" ./your-ruby-project

# 匹配Regexp.new构造的正则
grep -rE 'Regexp\.new\(["'\''"].*["'\'']\)' --include="*.rb" ./your-ruby-project

⚠️ 局限性:无法区分字符串中的/和真实正则,也无法处理动态生成的正则(如变量拼接的模式)。


内容的提问来源于stack exchange,提问作者Kakashi77

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 12:45:37