如何从带随机前后缀的物品名称提取对应ID(Python实现)
批量处理带随机前后缀/特殊符号的物品名称映射ID方案
针对需要处理3000+带随机前后缀、Unicode符号的物品名称,匹配对应ID的需求,这里提供几个实用的Python实现方案:
方法一:核心关键词匹配(最稳定)
先把原始的名称-ID映射反向构建为ID到核心关键词的集合,再对输入名称做清洗后匹配关键词覆盖率最高的条目,适合前后缀无规律的场景。
# 原始常规名称-ID映射(替换为你的3000+条目) base_name_id = { "Helmet of Divan": "DIVAN_HELMET", "Sword of the Bear": "BEAR_SWORD", # ... 其他条目 } # 反向构建ID到核心关键词集合的映射 id_to_keywords = {} for name, item_id in base_name_id.items(): # 清洗核心名称:转小写、去掉冗余词、拆分单词为集合 cleaned_core = name.lower().replace("of", "").strip() core_words = set(cleaned_core.split()) id_to_keywords[item_id] = core_words def get_item_id(modified_name): # 清洗输入名称:转小写、移除特殊符号、拆分单词 special_chars = "✪!@#$%^&*()_+-=[]{}|;:,.<>?" cleaned_input = modified_name.lower().translate(str.maketrans("", "", special_chars)).strip() input_words = set(cleaned_input.split()) # 匹配关键词覆盖率最高的ID best_id = None max_match = 0 for item_id, core_words in id_to_keywords.items(): match_count = len(input_words & core_words) if match_count > max_match: max_match = match_count best_id = item_id return best_id # 测试示例 print(get_item_id("Wise Helmet of Divan")) # 输出: DIVAN_HELMET print(get_item_id("✪ Helmet of Divan ✪")) # 输出: DIVAN_HELMET print(get_item_id("Clean Helmet of Divan")) # 输出: DIVAN_HELMET
方法二:正则提取核心结构(针对固定格式名称)
如果物品名称大多遵循「前缀+核心名称(X of Y)+后缀/符号」的格式,用正则直接提取核心名称片段,再匹配原始映射。
import re base_name_id = { "Helmet of Divan": "DIVAN_HELMET", "Sword of the Bear": "BEAR_SWORD", # ... 其他条目 } # 预编译正则:匹配"[名词] of [名词]"的核心结构,忽略大小写 core_pattern = re.compile(r'\b(\w+ of \w+)\b', re.IGNORECASE) def get_item_id_regex(modified_name): # 提取核心名称片段 match_result = core_pattern.search(modified_name) if match_result: core_name = match_result.group(1).strip().title() # 转为首字母大写匹配原始字典 return base_name_id.get(core_name) # 正则匹配失败时,用关键词匹配兜底 special_chars = "✪!@#$%^&*()_+-=[]{}|;:,.<>?" cleaned_input = modified_name.lower().translate(str.maketrans("", "", special_chars)).strip() input_words = set(cleaned_input.split()) best_id = None max_match = 0 for name, item_id in base_name_id.items(): core_words = set(name.lower().split()) match_count = len(input_words & core_words) if match_count > max_match: max_match = match_count best_id = item_id return best_id # 测试示例 print(get_item_id_regex("Wise Helmet of Divan")) # 输出: DIVAN_HELMET print(get_item_id_regex("✪ Helmet of Divan ✪")) # 输出: DIVAN_HELMET
方法三:模糊匹配(处理极端不规则场景)
如果前面两种方法都有误差,可以用模糊匹配计算输入名称与原始名称的相似度,取最接近的结果(需要安装第三方库)。
from fuzzywuzzy import fuzz # 安装命令: pip install fuzzywuzzy python-Levenshtein base_name_id = { "Helmet of Divan": "DIVAN_HELMET", "Sword of the Bear": "BEAR_SWORD", # ... 其他条目 } def get_item_id_fuzzy(modified_name): # 清洗输入名称 special_chars = "✪!@#$%^&*()_+-=[]{}|;:,.<>?" cleaned_input = modified_name.lower().translate(str.maketrans("", "", special_chars)).strip() best_id = None highest_score = 0 threshold = 70 # 相似度阈值,低于则返回None for name, item_id in base_name_id.items(): cleaned_name = name.lower().strip() # 使用部分匹配分数,适配前后缀场景 score = fuzz.partial_ratio(cleaned_input, cleaned_name) if score > highest_score: highest_score = score best_id = item_id return best_id if highest_score >= threshold else None # 测试示例 print(get_item_id_fuzzy("Wise Helmet of Divan")) # 输出: DIVAN_HELMET
优化建议
- 先对3000+原始条目做去重处理,避免重复匹配
- 可以把清洗特殊字符的逻辑抽成单独函数,方便扩展需要移除的符号
- 如果匹配结果有歧义,可以输出相似度最高的前3个结果人工校验
内容的提问来源于stack exchange,提问作者Yellow Blood
相关产品推荐
相关产品推荐

