You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Sentence Transformer:如何按索引/更新日期排序获取匹配句子

问题描述

我有一个包含大量记录的数据库表,字段为id、sentence、info、updated_date,数据示例如下:

idsentenceinfoupdated_info_date
1What is the name of your companysome distinct info19/12/2022
2Company Namesome distinct info18/12/2022
3What is the name of your companysome distinct info17/12/2022
4What is the name of your companysome distinct info16/12/2022
5What is the name of your companysome distinct info15/12/2022
6What is the name of your companysome distinct info14/12/2022
7What is the name of your companysome distinct info13/12/2022
8What is the phone number of your companysome distinct info12/12/2022
9What is the name of your companysome distinct info11/12/2022
10What is the name of your companysome distinct info10/12/2022

我将这些sentence转换为张量后,用示例句子"What is the name of your company"的张量做匹配,代码如下:

sentence = "What is the name of your company" # 已转为张量格式
cos_scores = util.pytorch_cos_sim(sentence, all_sentences_tensors)[0]

top_results = torch.topk(cos_scores, k=5) 
# 或者用numpy的方式
top_results = np.argpartition(cos_scores, range(5))[0:5]

由于多个句子内容完全相同,相似度得分均为1,top_results返回的结果顺序随机,无法按索引或updated_date从新到旧排序。我想获取按最新updated_date或索引顺序排列的top5匹配结果,该怎么实现?

解决方案

完全可行,以下是几种实用实现思路:

思路1:先筛选匹配项,再二次排序

  1. 先找出所有相似度得分等于1的记录索引(即完全匹配的句子)
  2. 根据这些索引从数据集中取出对应的id和updated_date
  3. 按updated_date降序(最新优先)或id降序排序,再取前5条

代码示例:

import pandas as pd
import torch

# 获取所有得分等于1的索引(cos_scores为torch张量)
match_indices = torch.where(cos_scores == 1.0)[0].cpu().numpy()

# 假设df是加载好的数据框,筛选匹配记录
matched_df = df.loc[match_indices]

# 按updated_date降序取前5
sorted_df = matched_df.sort_values(by='updated_info_date', ascending=False).head(5)
# 若按id降序则用:
# sorted_df = matched_df.sort_values(by='id', ascending=False).head(5)

# 最终得到有序的top5索引
top_sorted_indices = sorted_df.index.tolist()

思路2:融合日期/索引权重,一次排序出结果

把updated_date转为数值(如时间戳)或直接用id,和相似度得分结合成综合排序依据,一次排序即可得到结果:

import pandas as pd
import torch

# 将日期转换为时间戳并归一化(确保权重不盖过相似度得分)
df['timestamp'] = pd.to_datetime(df['updated_info_date'], format='%d/%m/%Y').astype('int64') / 10**9
norm_timestamp = torch.tensor((df['timestamp'] - df['timestamp'].min()) / (df['timestamp'].max() - df['timestamp'].min())).float()

# 构建综合得分:相似度得分 + 归一化时间戳*权重(权重可根据需求调整)
combined_scores = cos_scores + 0.1 * norm_timestamp

# 按综合得分取top5,得分相同时会按时间戳排序
top_results = torch.topk(combined_scores, k=5)
# top_results.indices就是最终有序的索引列表

思路3:直接用数据库查询(高效首选)

如果数据存在数据库中,直接用SQL筛选+排序效率最高,无需加载全量数据做张量匹配:

SELECT * FROM your_table 
WHERE sentence = 'What is the name of your company' 
ORDER BY updated_info_date DESC 
LIMIT 5;

完全匹配的场景下,字符串比对比张量相似度计算更高效,还能直接实现排序需求。

内容的提问来源于stack exchange,提问作者Avi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 14:01:43