Python Sentence Transformer:如何按索引/更新日期排序获取匹配句子
问题描述
我有一个包含大量记录的数据库表,字段为id、sentence、info、updated_date,数据示例如下:
| id | sentence | info | updated_info_date |
|---|---|---|---|
| 1 | What is the name of your company | some distinct info | 19/12/2022 |
| 2 | Company Name | some distinct info | 18/12/2022 |
| 3 | What is the name of your company | some distinct info | 17/12/2022 |
| 4 | What is the name of your company | some distinct info | 16/12/2022 |
| 5 | What is the name of your company | some distinct info | 15/12/2022 |
| 6 | What is the name of your company | some distinct info | 14/12/2022 |
| 7 | What is the name of your company | some distinct info | 13/12/2022 |
| 8 | What is the phone number of your company | some distinct info | 12/12/2022 |
| 9 | What is the name of your company | some distinct info | 11/12/2022 |
| 10 | What is the name of your company | some distinct info | 10/12/2022 |
我将这些sentence转换为张量后,用示例句子"What is the name of your company"的张量做匹配,代码如下:
sentence = "What is the name of your company" # 已转为张量格式 cos_scores = util.pytorch_cos_sim(sentence, all_sentences_tensors)[0] top_results = torch.topk(cos_scores, k=5) # 或者用numpy的方式 top_results = np.argpartition(cos_scores, range(5))[0:5]
由于多个句子内容完全相同,相似度得分均为1,top_results返回的结果顺序随机,无法按索引或updated_date从新到旧排序。我想获取按最新updated_date或索引顺序排列的top5匹配结果,该怎么实现?
解决方案
完全可行,以下是几种实用实现思路:
思路1:先筛选匹配项,再二次排序
- 先找出所有相似度得分等于1的记录索引(即完全匹配的句子)
- 根据这些索引从数据集中取出对应的
id和updated_date - 按
updated_date降序(最新优先)或id降序排序,再取前5条
代码示例:
import pandas as pd import torch # 获取所有得分等于1的索引(cos_scores为torch张量) match_indices = torch.where(cos_scores == 1.0)[0].cpu().numpy() # 假设df是加载好的数据框,筛选匹配记录 matched_df = df.loc[match_indices] # 按updated_date降序取前5 sorted_df = matched_df.sort_values(by='updated_info_date', ascending=False).head(5) # 若按id降序则用: # sorted_df = matched_df.sort_values(by='id', ascending=False).head(5) # 最终得到有序的top5索引 top_sorted_indices = sorted_df.index.tolist()
思路2:融合日期/索引权重,一次排序出结果
把updated_date转为数值(如时间戳)或直接用id,和相似度得分结合成综合排序依据,一次排序即可得到结果:
import pandas as pd import torch # 将日期转换为时间戳并归一化(确保权重不盖过相似度得分) df['timestamp'] = pd.to_datetime(df['updated_info_date'], format='%d/%m/%Y').astype('int64') / 10**9 norm_timestamp = torch.tensor((df['timestamp'] - df['timestamp'].min()) / (df['timestamp'].max() - df['timestamp'].min())).float() # 构建综合得分:相似度得分 + 归一化时间戳*权重(权重可根据需求调整) combined_scores = cos_scores + 0.1 * norm_timestamp # 按综合得分取top5,得分相同时会按时间戳排序 top_results = torch.topk(combined_scores, k=5) # top_results.indices就是最终有序的索引列表
思路3:直接用数据库查询(高效首选)
如果数据存在数据库中,直接用SQL筛选+排序效率最高,无需加载全量数据做张量匹配:
SELECT * FROM your_table WHERE sentence = 'What is the name of your company' ORDER BY updated_info_date DESC LIMIT 5;
完全匹配的场景下,字符串比对比张量相似度计算更高效,还能直接实现排序需求。
内容的提问来源于stack exchange,提问作者Avi
相关产品推荐
相关产品推荐

