You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言实现按ID统计动物词频并生成独热计数列

解决R语言中动物频次统计并转宽表的问题

嗨,我来帮你搞定这个需求!咱们可以用tidyverse工具包来轻松实现,步骤清晰又直观,下面是具体的操作方法:

步骤1:加载必要的包

首先确保你已经安装了tidyverse,如果没安装的话先运行install.packages("tidyverse"),然后加载它:

library(tidyverse)

步骤2:导入你的示例数据

先把你提供的示例数据加载进来:

df <- structure(list(id = 1:3, animals = structure(1:3, .Label = c("cat dog", "cat pig dog fish fish", "horse horse"), class = "factor")), .Names = c("id", "animals"), class = "data.frame", row.names = c(NA, -3L))

步骤3:数据处理与转换

通过三步操作得到你想要的结果:

df_result <- df %>%
  # 1. 把animals列按空格拆分,每个动物单独占一行
  separate_rows(animals, sep = " ") %>%
  # 2. 按id和动物分组,统计每个id下每种动物的出现次数
  count(id, animals, name = "count") %>%
  # 3. 转换为宽格式,缺失的动物频次用0填充
  pivot_wider(names_from = animals, values_from = count, values_fill = 0)

查看结果

运行print(df_result)就能得到你想要的输出:

# A tibble: 3 × 6
     id   cat   dog  fish horse   pig
  <int> <int> <int> <int> <int> <int>
1     1     1     1     0     0     0
2     2     1     1     2     0     1
3     3     0     0     0     2     0

补充说明

  • separate_rows():专门用来把包含多个值的字符串拆分成多行,这里按空格分隔正好匹配你的数据格式;
  • count():替代group_by() %>% summarize(count = n())的快捷方式,直接统计分组后的数量;
  • pivot_wider():把长格式数据转换为宽格式,values_fill = 0确保那些在某个id下没出现的动物显示为0,而不是NA。

内容的提问来源于stack exchange,提问作者0mm3

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:56:28