You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在dplyr的distinct()中用列名向量实现多列去重?

Great question! I’ve run into this exact issue before with dplyr’s distinct functions, especially when dealing with wide datasets. Let’s break down what’s happening and how to fix it cleanly.

Why Your Original Approach Didn’t Work

The problem with your distinct_() attempts is that this older standard-evaluation function expects individual column arguments, not a single vector. When you pass exclude.columns directly, it only uses the first element of the vector for deduplication—which is why swapping the order of columns in your vector gave different results (it was only deduplicating on the first column each time). The _() suffix functions like distinct_() are also deprecated in modern dplyr, so we have better tools now.

Single-Function Solution with Modern dplyr

Using across() and tidy evaluation, you can easily deduplicate based on all columns except your excluded subset (or a specific subset of columns) while retaining all columns in your output. Here’s how:

library(dplyr)
data(iris)

# Define columns to exclude from deduplication logic
exclude.columns <- c('Species', 'Sepal.Width')

# Deduplicate on all columns except the excluded ones, keep all original columns
deduplicated_iris <- iris %>%
  distinct(across(-all_of(exclude.columns)), .keep_all = TRUE)

Let’s break down the key parts:

  • across(-all_of(exclude.columns)): This tells dplyr to target every column except those in your exclude.columns vector. all_of() ensures the vector is correctly interpreted as column names.
  • .keep_all = TRUE: This critical argument keeps all columns (including your excluded ones) in the final output. Without it, distinct() would only return the columns used for deduplication.

If you ever want to deduplicate only a specific subset of columns (instead of excluding a few), you can adjust the syntax like this:

keep.columns <- c('Sepal.Length', 'Petal.Length', 'Petal.Width')
iris %>%
  distinct(across(all_of(keep.columns)), .keep_all = TRUE)

This approach eliminates the need to subset, deduplicate, and merge your data back together—keeping your code concise and easy to read.

内容的提问来源于stack exchange,提问作者Ollinator

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:33:28