You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从DataFrame指定列中提取每类值的唯一去重样本

获取DataFrame指定列去重唯一值的实现方案

以下是R语言环境下的两种常用实现方式,适配你提到的gapminder数据集场景:

基础R语法实现

直接调用内置的unique()函数即可完成列去重:

  • 提取国家列唯一值:
    代码:unique(gapminder$country)
    若需要逐行打印输出,使用:cat(unique(gapminder$country), sep = "\n")
  • 提取大洲列唯一值:
    代码:unique(gapminder$continent)
    逐行打印输出:cat(unique(gapminder$continent), sep = "\n")

运行大洲列代码的输出示例:

Africa
Antarctica
Asia
Australia
Europe
North America
South America

tidyverse生态实现(适配tibble格式)

如果使用dplyr包处理,可结合distinct()和pull()函数实现:

  • 提取国家列唯一值:
    代码:
    library(dplyr)
    gapminder %>% 
      distinct(country) %>% 
      pull(country)
    
    逐行打印的写法:
    gapminder %>% 
      distinct(country) %>% 
      pull(country) %>% 
      cat(sep = "\n")
    
  • 额外需求:如果需要保留每个唯一值对应的其他列样本记录,添加.keep_all = TRUE参数即可:
    代码:gapminder %>% distinct(country, .keep_all = TRUE)
    该代码会返回每个国家对应的任意一条完整样本行,包含lifeExp、pop等其他字段数据。

内容的提问来源于stack exchange,提问作者Eugene Vlasov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 11:45:04