如何利用BERT与余弦相似度实现电商商品名称标准化?
商品名称标准化解决方案
问题背景
我从不同超市网站采集商品数据时,发现同一款商品在各平台的名称表述差异很大,比如:
| 商品名称 | 超市 |
|---|---|
| 360° Toothbrush With Tongue And Cheek Cleaner | NoFrills |
| MEDIUM TOOTHBRUSH, 360° | FoodBasics |
我希望将这类名称统一标准化为「360° Toothbrush」。目前已经实现了用BERT提取句向量并计算余弦相似度的方法(能区分相似商品,相似度数值更高),但不知道如何利用这些相似度结果生成标准化名称后存入数据库。
测试脚本如下:
import torch from transformers import BertTokenizer, BertModel def get_sentence_vector(text): # Initialize the tokenizer and model tokenizer = BertTokenizer.from_pretrained('bert-base-uncased') model = BertModel.from_pretrained('bert-base-uncased') # Ensure the model is in evaluation mode model.eval() # Encode the text inputs = tokenizer(text, return_tensors='pt') # Perform a forward pass without gradient calculation with torch.no_grad(): outputs = model(**inputs) # Extract the last hidden state last_hidden_states = outputs.last_hidden_state # Aggregate the hidden states using mean pooling sentence_vector = torch.mean(last_hidden_states, dim=1) return sentence_vector def cosine_similarity(vec1, vec2): # Calculate the cosine similarity cosine_sim = torch.nn.functional.cosine_similarity(vec1, vec2, dim=1) return cosine_sim.item() example_text = "360° Toothbrush With Tongue And Cheek Cleaner" example_text_2 = "MEDIUM TOOTHBRUSH, 360°" example_text_3 = "Hair Expertise Hyaluron Plump Shampoo, with Hyaluronic Acid" vector = get_sentence_vector(example_text) vector_2 = get_sentence_vector(example_text_2) vector_3 = get_sentence_vector(example_text_3) print(cosine_similarity(vector, vector_2)) # 0.8426903486251831 print(cosine_similarity(vector, vector_3)) # 0.7190334796905518 print(cosine_similarity(vector_2, vector_3)) # 0.622465193271637
解决方案
1. 先优化现有代码
你的脚本每次调用get_sentence_vector都重新加载BERT模型,效率极低,先把模型初始化移到函数外面:
import torch from transformers import BertTokenizer, BertModel # 全局初始化模型和分词器,避免重复加载 tokenizer = BertTokenizer.from_pretrained('bert-base-uncased') model = BertModel.from_pretrained('bert-base-uncased') model.eval() def get_sentence_vector(text): inputs = tokenizer(text, return_tensors='pt', padding=True, truncation=True) with torch.no_grad(): outputs = model(**inputs) last_hidden_states = outputs.last_hidden_state # 用[CLS] token的向量更稳定,或者继续用均值池化 sentence_vector = last_hidden_states[:, 0, :] # 取[CLS]向量 # sentence_vector = torch.mean(last_hidden_states, dim=1) # 原均值池化 return sentence_vector def cosine_similarity(vec1, vec2): return torch.nn.functional.cosine_similarity(vec1, vec2, dim=1).item()
2. 基于相似度的标准化流程
步骤1:聚类分组,把相似商品归为一类
用余弦相似度做聚类,把相似度高于阈值(比如0.8,可根据测试调整)的商品名称分到同一组:
- 遍历所有采集到的商品名称,为每个名称生成向量
- 用DBSCAN或层次聚类算法,根据向量间的相似度分组(DBSCAN适合自动发现类别,无需预设聚类数)
- 示例伪代码:
from sklearn.cluster import DBSCAN import numpy as np # 假设所有商品名称存在列表all_product_names中 all_vectors = [get_sentence_vector(name).numpy().flatten() for name in all_product_names] X = np.array(all_vectors) # DBSCAN聚类,eps是相似度转化的距离(1-相似度),min_samples是最小聚类大小 dbscan = DBSCAN(eps=0.2, min_samples=2, metric='cosine') clusters = dbscan.fit_predict(X) # 把同一聚类的商品名称分组 cluster_groups = {} for idx, label in enumerate(clusters): if label not in cluster_groups: cluster_groups[label] = [] cluster_groups[label].append(all_product_names[idx])
步骤2:为每个聚类生成标准化名称
对每个聚类组,可通过以下方式生成标准名:
- 人工预设模板:提前整理常见商品的标准名模板,匹配组内关键词(比如检测到"360°"和"Toothbrush",直接用预设的「360° Toothbrush」)
- 自动提取核心关键词:用TF-IDF或RAKE工具从组内名称中提取最核心的关键词组合
- 选最简洁名称:从组内挑选长度最短、不含冗余描述(如With Tongue And Cheek Cleaner、MEDIUM)的名称
- 投票式生成:统计组内出现频率最高的词汇,拼接成标准名
示例代码(核心关键词提取+模板匹配):
import re from collections import Counter def get_standard_name(group, template_map=None): # 统一小写并去除标点 cleaned_names = [re.sub(r'[^\w\s°]', '', name.lower()) for name in group] # 拆分所有词汇 all_words = [] for name in cleaned_names: all_words.extend(name.split()) # 过滤冗余词并统计词频 stop_words = {'with', 'and', 'medium', 'cleaner', 'cheek', 'tongue'} word_counts = Counter([word for word in all_words if word not in stop_words]) # 取Top2核心词 core_words = [word for word, _ in word_counts.most_common(2)] # 按格式排序(数字符号在前,商品名在后) core_words.sort(key=lambda x: (not bool(re.search(r'^\d+°', x)), x)) standard_name = ' '.join(core_words).title() # 优先匹配预设模板 if template_map: for key in template_map: if all(kw in standard_name.lower() for kw in key.lower().split()): return template_map[key] return standard_name # 预设模板示例 template_map = { '360° toothbrush': '360° Toothbrush', 'hyaluron plump shampoo': 'Hair Expertise Hyaluron Plump Shampoo' } # 测试分组 test_group = ["360° Toothbrush With Tongue And Cheek Cleaner", "MEDIUM TOOTHBRUSH, 360°"] print(get_standard_name(test_group, template_map)) # 输出:360° Toothbrush
步骤3:批量映射并存入数据库
- 为每个原始商品名称匹配对应的标准名
- 将原始名称、标准名、超市信息等字段一起存入数据库
- 后续采集新商品时,先计算其与现有标准名向量的相似度,若高于阈值则直接复用对应标准名,否则新增标准名
3. 进阶优化
- 结合商品属性辅助判断:比如商品分类(oral care)、价格、品牌等信息,提升聚类准确性
- 微调BERT模型:用自有商品名称数据微调模型,让它更擅长区分同品类不同商品的名称
- 建立标准名字典:积累数据后构建标准名-原始名映射字典,后续直接匹配关键词,效率更高
内容的提问来源于stack exchange,提问作者eulerisgood
相关产品推荐
相关产品推荐

