如何在Azure Cognitive Search中为Collections(Edm.ComplexType)字段添加Freshness评分函数?
解决方案步骤
要实现按最新TimeSlot排序搜索结果,核心是给每个文档新增一个存储最新Slot时间的单个DateTimeOffset字段,再基于该字段配置Freshness评分函数,具体步骤如下:
1. 修改索引结构,新增单个时间字段
在现有索引中添加一个类型为Edm.DateTimeOffset的字段(比如命名为LatestTimeSlot),并确保该字段开启以下属性:
filterable: truesortable: truesearchable: false(可选,无需搜索该字段时可关闭)retrievable: true(可选,需返回给前端时开启)
示例索引字段定义片段:
{ "name": "LatestTimeSlot", "type": "Edm.DateTimeOffset", "filterable": true, "sortable": true, "searchable": false }
2. 填充LatestTimeSlot字段
有两种可靠方式填充该字段,可根据数据导入流程选择:
方式一:数据预处理(推荐,性能更优)
在将文档推送至Azure Cognitive Search前,自行计算每个文档TimeSlots数组中的最大Slot时间,赋值给LatestTimeSlot字段。
比如Python处理逻辑示例:
def get_latest_slot(time_slots): if not time_slots: return None slots = [slot["Slot"] for slot in time_slots] return max(slots) # 处理单个文档 document = { "TimeSlots": [ {"Slot": "2020-11-23T08:00:00-08:00"}, {"Slot": "2020-11-23T09:00:00-08:00"}, {"Slot": "2023-11-23T10:00:00-08:00"} ] } document["LatestTimeSlot"] = get_latest_slot(document["TimeSlots"])
方式二:使用Azure Cognitive Search技能集(适合无法修改数据源的场景)
若无法在数据源端预处理,可通过搜索服务的自定义技能计算最新Slot时间:
- 创建自定义技能(如用Azure Function实现),接收
TimeSlots数组,遍历返回最大Slot值。 - 在索引器的技能集中添加该自定义技能,将输出映射到
LatestTimeSlot字段。 - 重新运行索引器完成字段填充。
3. 配置Freshness评分函数
在搜索配置中,基于LatestTimeSlot字段创建Freshness评分函数,设置合适的衰减参数(如衰减周期、衰减量),让最新时间的文档获得更高评分,实现排序效果。
示例评分配置片段:
{ "scoringProfiles": [ { "name": "FreshnessScoring", "functions": [ { "type": "freshness", "fieldName": "LatestTimeSlot", "boost": 5, "freshness": { "boostingDuration": "P365D" // 1年内的文档获得衰减提升 } } ], "functionAggregation": "sum" } ] }
之后在搜索请求中指定使用该评分配置,或直接按LatestTimeSlot降序排序,即可实现按最新Slot时间排序的需求。
内容的提问来源于stack exchange,提问作者Vadiraj Rao
相关产品推荐
相关产品推荐

