Lua使用luastring.translate时子串匹配冲突导致翻译错误如何解决
魔方公式转译匹配冲突修复方案
问题背景
开发魔方转动公式转译工具,目标是将标准公式序列R U R' U'转译为自定义格式r^ u< rv u>。实际调用luastring库的translate接口执行替换时出现匹配逻辑错误:带撇号的转动标记(如R')包含普通短标记子串R,接口默认短串优先匹配,输入R U R' U'时会错误输出r^ u< r^' u<',撇号残留,结果不符合预期。
此前已尝试的实现思路:
- 编写自定义
split字符串分割函数 - 定义基础转动标记数组
moves与对应目标标记数组neomoves - 构建
translation_table翻译映射表
但始终无法实现长标记(带'的转动符号)优先匹配,luastring库未提供相关匹配优先级配置,公开技术社区也未找到对应解法,原有测试代码如下:
io.write("put in your algorithm: ") local alg = io.read() function split(str, pat) local t = {} local fpat = "(.-)" .. pat local last_end = 1 local s, e, cap = str:find(fpat, 1) while s do if s ~= 1 or cap ~= "" then table.insert(t, cap) end last_end = e + 1 s, e, cap = str:find(fpat, last_end) end if last_end <= #str then cap = str:sub(last_end) table.insert(t, cap) end return t end -- table stuff so we can see local moves = {"R", "U", "L", "F", "D", "B", "M", "E", "S", "R'", "U'", "L'", "F'", "D'", "B'", "M'", "E'", "S'"} local neomoves = {"r^", "u<", "lv", "f>", "d>", "b<", "mv", "e>", "s>", "rv", "u>", "l^", "f<", "d<", "b>", "m^", "e<", "s<"} local string = require("luastring") local translation_table = { ["R'"] = "rv", ["U'"] = "u<", ["L'"] = "lv", ["F'"] = "f>", ["D'"] = "d>", ["B'"] = "b<", ["M'"] = "mv", ["E'"] = "e>", ["S'"] = "s>", ["R"] = "r^", ["U"] = "u<", ["L"] = "lv", ["F"] = "f>", ["D"] = "d>", ["B"] = "b<", ["M"] = "mv", ["E"] = "e>", ["S"] = "s>" } -- translate the moves to neomoves local translated = string.translate(alg, translation_table) print(translated)
注:原有代码的映射表本身存在笔误,
U'对应的目标值按照neomoves数组定义应为u>,原代码错写为u<。
故障原因
luastring库的translate接口采用逐字符顺序替换逻辑,没有长串优先匹配机制,遍历到R'的首字符R时会直接命中短标记R的替换规则,不会向后检查是否存在更长的合法标记,最终导致撇号残留、转译错误。
修复方案
放弃依赖luastring的translate接口,手动实现逐位扫描+长标记优先匹配逻辑,不需要复杂的字符串分割:
- 将所有合法转动标记按长度从大到小排序,保证2字符的带撇标记排在1字符普通标记前面
- 逐位遍历输入字符串,跳过空格分隔符,每次从当前位置优先尝试匹配最长的合法标记
- 匹配成功后追加对应转译结果,指针向后移动对应标记的长度,直到遍历完成整个输入
修复后的可运行代码如下:
io.write("put in your algorithm: ") local alg = io.read() -- 对齐moves与neomoves的正确映射关系 local translation_table = { ["R'"] = "rv", ["U'"] = "u>", ["L'"] = "l^", ["F'"] = "f<", ["D'"] = "d<", ["B'"] = "b>", ["M'"] = "m^", ["E'"] = "e<", ["S'"] = "s<", ["R"] = "r^", ["U"] = "u<", ["L"] = "lv", ["F"] = "f>", ["D"] = "d>", ["B"] = "b<", ["M"] = "mv", ["E"] = "e>", ["S"] = "s>" } -- 所有标记按长度降序排列,确保长标记优先匹配 local sorted_moves = {} for move, _ in pairs(translation_table) do table.insert(sorted_moves, move) end table.sort(sorted_moves, function(a, b) return #a > #b end) local result = {} local pos = 1 local alg_len = #alg while pos <= alg_len do -- 跳过空格分隔符 if alg:sub(pos, pos) == " " then pos = pos + 1 goto continue end -- 按优先级匹配标记 local is_matched = false for _, move in ipairs(sorted_moves) do local move_len = #move if pos + move_len - 1 <= alg_len and alg:sub(pos, pos + move_len - 1) == move then table.insert(result, translation_table[move]) pos = pos + move_len is_matched = true break end end -- 非法字符直接跳过,可按需改为抛出错误 if not is_matched then pos = pos + 1 end ::continue:: end -- 拼接结果输出 local translated = table.concat(result, " ") print(translated)
代码测试效果:输入R U R' U'时,将正确输出r^ u< rv u>,完全符合预期。后续如果需要新增更长的标记(如R2、r双层转、转体标记等),只需要在translation_table中添加对应键值对,排序逻辑会自动处理匹配优先级,不会再出现子串冲突问题。
内容的提问来源于stack exchange,提问作者Drogobo
相关产品推荐
相关产品推荐

