You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从带随机前后缀的物品名称提取对应ID(Python实现)

批量处理带随机前后缀/特殊符号的物品名称映射ID方案

针对需要处理3000+带随机前后缀、Unicode符号的物品名称,匹配对应ID的需求,这里提供几个实用的Python实现方案:

方法一:核心关键词匹配(最稳定)

先把原始的名称-ID映射反向构建为ID到核心关键词的集合,再对输入名称做清洗后匹配关键词覆盖率最高的条目,适合前后缀无规律的场景。

# 原始常规名称-ID映射(替换为你的3000+条目)
base_name_id = {
    "Helmet of Divan": "DIVAN_HELMET",
    "Sword of the Bear": "BEAR_SWORD",
    # ... 其他条目
}

# 反向构建ID到核心关键词集合的映射
id_to_keywords = {}
for name, item_id in base_name_id.items():
    # 清洗核心名称:转小写、去掉冗余词、拆分单词为集合
    cleaned_core = name.lower().replace("of", "").strip()
    core_words = set(cleaned_core.split())
    id_to_keywords[item_id] = core_words

def get_item_id(modified_name):
    # 清洗输入名称:转小写、移除特殊符号、拆分单词
    special_chars = "✪!@#$%^&*()_+-=[]{}|;:,.<>?"
    cleaned_input = modified_name.lower().translate(str.maketrans("", "", special_chars)).strip()
    input_words = set(cleaned_input.split())
    
    # 匹配关键词覆盖率最高的ID
    best_id = None
    max_match = 0
    for item_id, core_words in id_to_keywords.items():
        match_count = len(input_words & core_words)
        if match_count > max_match:
            max_match = match_count
            best_id = item_id
    return best_id

# 测试示例
print(get_item_id("Wise Helmet of Divan"))  # 输出: DIVAN_HELMET
print(get_item_id("✪ Helmet of Divan ✪"))  # 输出: DIVAN_HELMET
print(get_item_id("Clean Helmet of Divan"))  # 输出: DIVAN_HELMET

方法二:正则提取核心结构(针对固定格式名称)

如果物品名称大多遵循「前缀+核心名称(X of Y)+后缀/符号」的格式,用正则直接提取核心名称片段,再匹配原始映射。

import re

base_name_id = {
    "Helmet of Divan": "DIVAN_HELMET",
    "Sword of the Bear": "BEAR_SWORD",
    # ... 其他条目
}

# 预编译正则:匹配"[名词] of [名词]"的核心结构,忽略大小写
core_pattern = re.compile(r'\b(\w+ of \w+)\b', re.IGNORECASE)

def get_item_id_regex(modified_name):
    # 提取核心名称片段
    match_result = core_pattern.search(modified_name)
    if match_result:
        core_name = match_result.group(1).strip().title()  # 转为首字母大写匹配原始字典
        return base_name_id.get(core_name)
    
    # 正则匹配失败时,用关键词匹配兜底
    special_chars = "✪!@#$%^&*()_+-=[]{}|;:,.<>?"
    cleaned_input = modified_name.lower().translate(str.maketrans("", "", special_chars)).strip()
    input_words = set(cleaned_input.split())
    
    best_id = None
    max_match = 0
    for name, item_id in base_name_id.items():
        core_words = set(name.lower().split())
        match_count = len(input_words & core_words)
        if match_count > max_match:
            max_match = match_count
            best_id = item_id
    return best_id

# 测试示例
print(get_item_id_regex("Wise Helmet of Divan"))  # 输出: DIVAN_HELMET
print(get_item_id_regex("✪ Helmet of Divan ✪"))  # 输出: DIVAN_HELMET

方法三:模糊匹配(处理极端不规则场景)

如果前面两种方法都有误差,可以用模糊匹配计算输入名称与原始名称的相似度,取最接近的结果(需要安装第三方库)。

from fuzzywuzzy import fuzz  # 安装命令: pip install fuzzywuzzy python-Levenshtein

base_name_id = {
    "Helmet of Divan": "DIVAN_HELMET",
    "Sword of the Bear": "BEAR_SWORD",
    # ... 其他条目
}

def get_item_id_fuzzy(modified_name):
    # 清洗输入名称
    special_chars = "✪!@#$%^&*()_+-=[]{}|;:,.<>?"
    cleaned_input = modified_name.lower().translate(str.maketrans("", "", special_chars)).strip()
    
    best_id = None
    highest_score = 0
    threshold = 70  # 相似度阈值,低于则返回None
    
    for name, item_id in base_name_id.items():
        cleaned_name = name.lower().strip()
        # 使用部分匹配分数,适配前后缀场景
        score = fuzz.partial_ratio(cleaned_input, cleaned_name)
        if score > highest_score:
            highest_score = score
            best_id = item_id
    
    return best_id if highest_score >= threshold else None

# 测试示例
print(get_item_id_fuzzy("Wise Helmet of Divan"))  # 输出: DIVAN_HELMET

优化建议

  1. 先对3000+原始条目做去重处理,避免重复匹配
  2. 可以把清洗特殊字符的逻辑抽成单独函数,方便扩展需要移除的符号
  3. 如果匹配结果有歧义,可以输出相似度最高的前3个结果人工校验

内容的提问来源于stack exchange,提问作者Yellow Blood

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 03:24:58