如何通过规范化算法或科学评分策略筛选多维度最优选项?
寻求多特征优先级筛选的规范化算法
我想找这类问题的通用算法或方法(需要知道对应的名称/关键词方便检索):我有一组有效选项,需要根据给定的优先级标准筛选出最优项。
以Atom feed条目里的多个<link>为例:
<entry> <title>Multiple Links</title> <link rel="self" type="text/plain" href="http://example.org/2003/12/13/atom03.txt"/> <link rel="alternate" type="text/html" href="http://example.org/2003/12/13/atom03.html"/> <link rel="alternate" type="application/json" href="http://example.org/2003/12/13/atom03.json"/> </entry>
这种情况下很难直接确定该选哪个链接。
我的目标是找到可在浏览器中阅读文章的链接:根据规范,这类链接应标记为rel="alternate";备选是条目源文件(比如生成feed的markdown文件),对应rel="self"。由于要在浏览器中阅读,优先选择type="text/html",但浏览器可渲染的type="text/plain"等其他文本类型也可接受。
针对这个问题我已有可行解决方案:
defp pick_link(links) when is_list(links) do links |> Enum.sort_by(fn {_link, rel, type} -> score_rel(rel) + score_type(type) end, :desc) |> List.first() end defp score_rel("alternate"), do: 5 defp score_rel("self"), do: 2 defp score_rel(_rel), do: 0 defp score_type("text/html"), do: 3 defp score_type("text/" <> _rest), do: 1 defp score_type(_type), do: 0
我通过给更优特征赋予更高分数,将得分相加后排序,取最高分作为最优选项。
但这个方法存在问题:评分是孤立计算的,例如rel="self" type="text/html"的得分低于rel="alternate" type="application/json",但前者其实更符合需求。我不确定怎么设置分数才能让特征组合的评分符合预期,而且特征数量越多,这个问题越复杂。
请问是否存在此类问题的规范化算法?或是更具数学合理性的评分赋值方式?
内容的提问来源于stack exchange,提问作者Lukas Knuth
相关产品推荐
相关产品推荐

