You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用BERT与余弦相似度实现电商商品名称标准化?

商品名称标准化解决方案

问题背景

我从不同超市网站采集商品数据时,发现同一款商品在各平台的名称表述差异很大,比如:

商品名称超市
360° Toothbrush With Tongue And Cheek CleanerNoFrills
MEDIUM TOOTHBRUSH, 360°FoodBasics

我希望将这类名称统一标准化为「360° Toothbrush」。目前已经实现了用BERT提取句向量并计算余弦相似度的方法(能区分相似商品,相似度数值更高),但不知道如何利用这些相似度结果生成标准化名称后存入数据库。

测试脚本如下:

import torch
from transformers import BertTokenizer, BertModel

def get_sentence_vector(text):
  
    # Initialize the tokenizer and model
    tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
    model = BertModel.from_pretrained('bert-base-uncased')

    # Ensure the model is in evaluation mode
    model.eval()

    # Encode the text
    inputs = tokenizer(text, return_tensors='pt')

    # Perform a forward pass without gradient calculation
    with torch.no_grad():
        outputs = model(**inputs)

    # Extract the last hidden state
    last_hidden_states = outputs.last_hidden_state

    # Aggregate the hidden states using mean pooling
    sentence_vector = torch.mean(last_hidden_states, dim=1)

    return sentence_vector


def cosine_similarity(vec1, vec2):
    # Calculate the cosine similarity
    cosine_sim = torch.nn.functional.cosine_similarity(vec1, vec2, dim=1)
    
    return cosine_sim.item() 


example_text = "360° Toothbrush With Tongue And Cheek Cleaner"
example_text_2 = "MEDIUM TOOTHBRUSH, 360°"
example_text_3 = "Hair Expertise Hyaluron Plump Shampoo, with Hyaluronic Acid"

vector = get_sentence_vector(example_text)
vector_2 = get_sentence_vector(example_text_2)
vector_3 = get_sentence_vector(example_text_3)

print(cosine_similarity(vector, vector_2)) # 0.8426903486251831
print(cosine_similarity(vector, vector_3)) # 0.7190334796905518
print(cosine_similarity(vector_2, vector_3)) # 0.622465193271637

解决方案

1. 先优化现有代码

你的脚本每次调用get_sentence_vector都重新加载BERT模型,效率极低,先把模型初始化移到函数外面:

import torch
from transformers import BertTokenizer, BertModel

# 全局初始化模型和分词器,避免重复加载
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
model = BertModel.from_pretrained('bert-base-uncased')
model.eval()

def get_sentence_vector(text):
    inputs = tokenizer(text, return_tensors='pt', padding=True, truncation=True)
    with torch.no_grad():
        outputs = model(**inputs)
    last_hidden_states = outputs.last_hidden_state
    # 用[CLS] token的向量更稳定,或者继续用均值池化
    sentence_vector = last_hidden_states[:, 0, :]  # 取[CLS]向量
    # sentence_vector = torch.mean(last_hidden_states, dim=1)  # 原均值池化
    return sentence_vector

def cosine_similarity(vec1, vec2):
    return torch.nn.functional.cosine_similarity(vec1, vec2, dim=1).item()

2. 基于相似度的标准化流程

步骤1:聚类分组,把相似商品归为一类

用余弦相似度做聚类,把相似度高于阈值(比如0.8,可根据测试调整)的商品名称分到同一组:

  • 遍历所有采集到的商品名称,为每个名称生成向量
  • 用DBSCAN或层次聚类算法,根据向量间的相似度分组(DBSCAN适合自动发现类别,无需预设聚类数)
  • 示例伪代码:
from sklearn.cluster import DBSCAN
import numpy as np

# 假设所有商品名称存在列表all_product_names中
all_vectors = [get_sentence_vector(name).numpy().flatten() for name in all_product_names]
X = np.array(all_vectors)

# DBSCAN聚类,eps是相似度转化的距离(1-相似度),min_samples是最小聚类大小
dbscan = DBSCAN(eps=0.2, min_samples=2, metric='cosine')
clusters = dbscan.fit_predict(X)

# 把同一聚类的商品名称分组
cluster_groups = {}
for idx, label in enumerate(clusters):
    if label not in cluster_groups:
        cluster_groups[label] = []
    cluster_groups[label].append(all_product_names[idx])

步骤2:为每个聚类生成标准化名称

对每个聚类组,可通过以下方式生成标准名:

  • 人工预设模板:提前整理常见商品的标准名模板,匹配组内关键词(比如检测到"360°"和"Toothbrush",直接用预设的「360° Toothbrush」)
  • 自动提取核心关键词:用TF-IDF或RAKE工具从组内名称中提取最核心的关键词组合
  • 选最简洁名称:从组内挑选长度最短、不含冗余描述(如With Tongue And Cheek Cleaner、MEDIUM)的名称
  • 投票式生成:统计组内出现频率最高的词汇,拼接成标准名

示例代码(核心关键词提取+模板匹配):

import re
from collections import Counter

def get_standard_name(group, template_map=None):
    # 统一小写并去除标点
    cleaned_names = [re.sub(r'[^\w\s°]', '', name.lower()) for name in group]
    # 拆分所有词汇
    all_words = []
    for name in cleaned_names:
        all_words.extend(name.split())
    # 过滤冗余词并统计词频
    stop_words = {'with', 'and', 'medium', 'cleaner', 'cheek', 'tongue'}
    word_counts = Counter([word for word in all_words if word not in stop_words])
    # 取Top2核心词
    core_words = [word for word, _ in word_counts.most_common(2)]
    # 按格式排序(数字符号在前,商品名在后)
    core_words.sort(key=lambda x: (not bool(re.search(r'^\d+°', x)), x))
    standard_name = ' '.join(core_words).title()
    # 优先匹配预设模板
    if template_map:
        for key in template_map:
            if all(kw in standard_name.lower() for kw in key.lower().split()):
                return template_map[key]
    return standard_name

# 预设模板示例
template_map = {
    '360° toothbrush': '360° Toothbrush',
    'hyaluron plump shampoo': 'Hair Expertise Hyaluron Plump Shampoo'
}

# 测试分组
test_group = ["360° Toothbrush With Tongue And Cheek Cleaner", "MEDIUM TOOTHBRUSH, 360°"]
print(get_standard_name(test_group, template_map))  # 输出:360° Toothbrush

步骤3:批量映射并存入数据库

  • 为每个原始商品名称匹配对应的标准名
  • 将原始名称、标准名、超市信息等字段一起存入数据库
  • 后续采集新商品时,先计算其与现有标准名向量的相似度,若高于阈值则直接复用对应标准名,否则新增标准名

3. 进阶优化

  • 结合商品属性辅助判断:比如商品分类(oral care)、价格、品牌等信息,提升聚类准确性
  • 微调BERT模型:用自有商品名称数据微调模型,让它更擅长区分同品类不同商品的名称
  • 建立标准名字典:积累数据后构建标准名-原始名映射字典,后续直接匹配关键词,效率更高

内容的提问来源于stack exchange,提问作者eulerisgood

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 06:55:57