如何用dplyr的slice_*函数改写带权重的top_n()调用?
用
slice_max()替代已弃用的top_n()实现分组取TOPN需求 我需要把代码中已弃用的top_n()调用替换为推荐的slice_max()函数,但不清楚如何在slice_max()中设置权重参数。
原代码(使用top_n()):
top10 <- structure( list( Variable = c("tfidf_text_crossing", "tfidf_text_best", "tfidf_text_amazing", "tfidf_text_fantastic", "tfidf_text_player", "tfidf_text_great", "tfidf_text_10", "tfidf_text_progress", "tfidf_text_relaxing", "tfidf_text_fix"), Importance = c(0.428820580430941, 0.412741988094224, 0.368676982306671, 0.361409225854695, 0.331176924533776, 0.307393456208119, 0.293945850296236, 0.286313554816565, 0.283457020779205, 0.27899280757397), Sign = c(tfidf_text_crossing = "POS", tfidf_text_best = "POS", tfidf_text_amazing = "POS", tfidf_text_fantastic = "POS", tfidf_text_player = "NEG", tfidf_text_great = "POS", tfidf_text_10 = "POS", tfidf_text_progress = "NEG", tfidf_text_relaxing = "POS", tfidf_text_fix = "NEG") ), row.names = c(NA, -10L), class = c("vi", "tbl_df", "tbl", "data.frame"), type = "|coefficient|" ) suppressPackageStartupMessages(library(dplyr)) top10 |> group_by(Sign) |> top_n(2, wt = abs(Importance))
运行结果:
# A tibble: 4 × 3 # Groups: Sign [2] Variable Importance Sign <chr> <dbl> <chr> 1 tfidf_text_crossing 0.429 POS 2 tfidf_text_best 0.413 POS 3 tfidf_text_player 0.331 NEG 4 tfidf_text_progress 0.286 NEG
我自己想到的替代写法是:
top10 |> group_by(Sign) |> arrange(desc(abs(Importance))) |> slice_head(n = 2)
但这种写法对新手来说可读性太差,有没有更直接的slice_*函数实现方式?
直接使用slice_max()的简洁实现
你可以直接用slice_max(),通过order_by参数指定排序依据(对应原top_n()的wt参数),n参数指定要选取的数量,一步到位完成分组取TOPN的需求,代码可读性对新手更友好:
top10 |> group_by(Sign) |> slice_max(order_by = abs(Importance), n = 2)
这段代码运行后会得到和原top_n()完全一致的结果,而且逻辑更直观:
slice_max()的函数名直接表明是选取最大值对应的行order_by明确指定了用来判断“最大”的字段(这里是abs(Importance)的绝对值)n清晰说明要取每组的前2个
为什么这个写法更适合教学?
相比arrange()+slice_head()的组合,slice_max()把排序和取值合并成了一步操作,新手不需要理解额外的排序逻辑,只需要知道“我要按某个字段取每组最大的n个”,就能直接对应到函数和参数,降低了理解门槛。
内容的提问来源于stack exchange,提问作者itsMeInMiami
相关产品推荐
相关产品推荐

