You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则提取Rails迁移文件中的add/remove_index命令?

解决Rails迁移文件中多行add/remove_index命令的正则提取问题

问题场景

需要从Rails迁移文件中提取指定表的add_index/remove_index命令并转为数组,单行命令场景下运行正常,但遇到多行书写的命令时,正则会将多个命令合并为一个字符串,无法正确拆分。

示例迁移文件:

class AddMissingUniqueIndices < ActiveRecord::Migration
  def self.up
    add_index :tags, :name, unique: true

    remove_index :taggings, :tag_id
    remove_index :taggings, [:taggable_id, :taggable_type, :context]
    add_index :taggings,
              [:tag_id, :taggable_id, :taggable_type, :context, :tagger_id, :tagger_type],
              unique: true, name: 'taggings_idx'
  end

  def self.down
    remove_index :tags, :name

    remove_index :taggings, name: 'taggings_idx'
    add_index :taggings, :tag_id
    add_index :taggings, [:taggable_id, :taggable_type, :context]
  end
end

当前错误代码及输出:

migration_content = 'migration file in txt'
@table_name = 'taggings'
regex_pattern = /(add|remove)_index\s+:#{@table_name}.*\w+:\s+?\w+/m
columns_to_process = migration_content.to_enum(:scan, regex_pattern).map { Regexp.last_match.to_s.squish }
puts columns_to_process
# => ["remove_index :taggings, :tag_id remove_index :taggings, [:taggable_id, :taggable_type, :context] add_index :taggings, [:tag_id, :taggable_id, :taggable_type, :context, :tagger_id, :tagger_type], unique: true"]

解决方案

核心思路

  1. 先提取self.up(或change)代码块,缩小匹配范围;
  2. 使用非贪婪匹配并明确命令结束边界,避免跨命令匹配;
  3. 准确匹配迁移命令的各种参数格式(数组、符号、字符串、键值对等)。

完整代码示例

migration_content = <<~MIGRATION
class AddMissingUniqueIndices < ActiveRecord::Migration
  def self.up
    add_index :tags, :name, unique: true

    remove_index :taggings, :tag_id
    remove_index :taggings, [:taggable_id, :taggable_type, :context]
    add_index :taggings,
              [:tag_id, :taggable_id, :taggable_type, :context, :tagger_id, :tagger_type],
              unique: true, name: 'taggings_idx'
  end

  def self.down
    remove_index :tags, :name

    remove_index :taggings, name: 'taggings_idx'
    add_index :taggings, :tag_id
    add_index :taggings, [:taggable_id, :taggable_type, :context]
  end
end
MIGRATION

# 提取self.up代码块,缩小匹配范围
up_block = migration_content.match(/def self.up\s+(.*?)\s+end/m)[1]

@table_name = 'taggings'
# 修正后的正则:匹配单个命令,明确结束边界
regex_pattern = /(add|remove)_index\s+:#{@table_name}(?:\s*,\s*(?:\[[^\]]+\]|:\w+|'.*?'|".*?"|\w+:\s*(?:\[[^\]]+\]|'.*?'|".*?"|\w+)))*?(?=\s*(?:(add|remove)_index|end|\z))/m

# 提取命令并格式化(去除多余空白)
columns_to_process = up_block.scan(regex_pattern).map do |match|
  (match[0] + $').split(/(?=(add|remove)_index)/)[0].squish
end

# 输出结果
p columns_to_process

输出结果

["remove_index :taggings, :tag_id", "remove_index :taggings, [:taggable_id, :taggable_type, :context]", "add_index :taggings, [:tag_id, :taggable_id, :taggable_type, :context, :tagger_id, :tagger_type], unique: true, name: 'taggings_idx'"]

正则说明

  • (add|remove)_index\s+:#{@table_name}:匹配命令开头及指定表名;
  • (?:\s*,\s*(?:\[[^\]]+\]|:\w+|'.*?'|".*?"|\w+:\s*(?:\[[^\]]+\]|'.*?'|".*?"|\w+)))*?:非贪婪匹配所有参数,支持数组、符号、字符串、键值对等格式;
  • (?=\s*(?:(add|remove)_index|end|\z)):正向预查命令结束边界,确保匹配到下一个命令、代码块结束或文本结尾时停止。

为什么原正则失效

原正则使用了贪婪的.*匹配,且未明确命令结束边界,导致会从第一个命令开始匹配到最后一个符合\w+:\s+?\w+的位置,直接将多个命令合并成了一个字符串。修正后的正则通过非贪婪匹配和边界限制,精准拆分每个独立命令。

内容的提问来源于stack exchange,提问作者dfop02

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 18:14:53