如何用正则提取Rails迁移文件中的add/remove_index命令?
解决Rails迁移文件中多行add/remove_index命令的正则提取问题
问题场景
需要从Rails迁移文件中提取指定表的add_index/remove_index命令并转为数组,单行命令场景下运行正常,但遇到多行书写的命令时,正则会将多个命令合并为一个字符串,无法正确拆分。
示例迁移文件:
class AddMissingUniqueIndices < ActiveRecord::Migration def self.up add_index :tags, :name, unique: true remove_index :taggings, :tag_id remove_index :taggings, [:taggable_id, :taggable_type, :context] add_index :taggings, [:tag_id, :taggable_id, :taggable_type, :context, :tagger_id, :tagger_type], unique: true, name: 'taggings_idx' end def self.down remove_index :tags, :name remove_index :taggings, name: 'taggings_idx' add_index :taggings, :tag_id add_index :taggings, [:taggable_id, :taggable_type, :context] end end
当前错误代码及输出:
migration_content = 'migration file in txt' @table_name = 'taggings' regex_pattern = /(add|remove)_index\s+:#{@table_name}.*\w+:\s+?\w+/m columns_to_process = migration_content.to_enum(:scan, regex_pattern).map { Regexp.last_match.to_s.squish } puts columns_to_process # => ["remove_index :taggings, :tag_id remove_index :taggings, [:taggable_id, :taggable_type, :context] add_index :taggings, [:tag_id, :taggable_id, :taggable_type, :context, :tagger_id, :tagger_type], unique: true"]
解决方案
核心思路
- 先提取
self.up(或change)代码块,缩小匹配范围; - 使用非贪婪匹配并明确命令结束边界,避免跨命令匹配;
- 准确匹配迁移命令的各种参数格式(数组、符号、字符串、键值对等)。
完整代码示例
migration_content = <<~MIGRATION class AddMissingUniqueIndices < ActiveRecord::Migration def self.up add_index :tags, :name, unique: true remove_index :taggings, :tag_id remove_index :taggings, [:taggable_id, :taggable_type, :context] add_index :taggings, [:tag_id, :taggable_id, :taggable_type, :context, :tagger_id, :tagger_type], unique: true, name: 'taggings_idx' end def self.down remove_index :tags, :name remove_index :taggings, name: 'taggings_idx' add_index :taggings, :tag_id add_index :taggings, [:taggable_id, :taggable_type, :context] end end MIGRATION # 提取self.up代码块,缩小匹配范围 up_block = migration_content.match(/def self.up\s+(.*?)\s+end/m)[1] @table_name = 'taggings' # 修正后的正则:匹配单个命令,明确结束边界 regex_pattern = /(add|remove)_index\s+:#{@table_name}(?:\s*,\s*(?:\[[^\]]+\]|:\w+|'.*?'|".*?"|\w+:\s*(?:\[[^\]]+\]|'.*?'|".*?"|\w+)))*?(?=\s*(?:(add|remove)_index|end|\z))/m # 提取命令并格式化(去除多余空白) columns_to_process = up_block.scan(regex_pattern).map do |match| (match[0] + $').split(/(?=(add|remove)_index)/)[0].squish end # 输出结果 p columns_to_process
输出结果
["remove_index :taggings, :tag_id", "remove_index :taggings, [:taggable_id, :taggable_type, :context]", "add_index :taggings, [:tag_id, :taggable_id, :taggable_type, :context, :tagger_id, :tagger_type], unique: true, name: 'taggings_idx'"]
正则说明
(add|remove)_index\s+:#{@table_name}:匹配命令开头及指定表名;(?:\s*,\s*(?:\[[^\]]+\]|:\w+|'.*?'|".*?"|\w+:\s*(?:\[[^\]]+\]|'.*?'|".*?"|\w+)))*?:非贪婪匹配所有参数,支持数组、符号、字符串、键值对等格式;(?=\s*(?:(add|remove)_index|end|\z)):正向预查命令结束边界,确保匹配到下一个命令、代码块结束或文本结尾时停止。
为什么原正则失效
原正则使用了贪婪的.*匹配,且未明确命令结束边界,导致会从第一个命令开始匹配到最后一个符合\w+:\s+?\w+的位置,直接将多个命令合并成了一个字符串。修正后的正则通过非贪婪匹配和边界限制,精准拆分每个独立命令。
内容的提问来源于stack exchange,提问作者dfop02
相关产品推荐
相关产品推荐

