You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

dplyr中mutate与transmute的区别及适用场景解析

mutate 与 transmute 的核心差异及适用场景

核心区别

  • mutate(): 保留数据框所有原有列,同时新增或修改指定列,最终输出列数 = 原列数 + 新增/修改列数(修改列不增加总数)。
  • transmute(): 默认只保留新增/修改的列,若要保留原列需明确指定,最终输出列数仅包含你明确指定的原列和新生成的列。

示例演示

先构造一个简单的数据框:

library(dplyr)
student_df <- tibble(
  student_id = 1:5,
  math_score = c(82, 95, 76, 88, 90)
)

使用 mutate()

保留原有所有列,同时添加新的计算列:

student_mutate <- student_df %>%
  mutate(
    math_score_boost = math_score + 5,
    grade = case_when(
      math_score >= 90 ~ "A",
      math_score >= 80 ~ "B",
      TRUE ~ "C"
    )
  )

print(student_mutate)
# 输出包含原列 student_id、math_score,以及新列 math_score_boost、grade

使用 transmute()

仅保留指定的原列和新生成的列,原列默认不保留:

student_transmute <- student_df %>%
  transmute(
    student_id,  # 明确保留原列 student_id
    math_score_double = math_score * 2,
    is_excellent = math_score >= 90
  )

print(student_transmute)
# 输出仅包含 student_id、math_score_double、is_excellent,原列 math_score 被丢弃

适用场景

  • mutate() 适用场景:
    • 数据清洗或预处理阶段,需要保留原始数据字段,同时衍生新字段用于后续分析(比如保留原始分数,同时计算排名、等级)。
    • 链式操作中,需要基于原列和新生成的列继续进行计算(比如先算总分,再算平均分)。
  • transmute() 适用场景:
    • 仅需要输出计算后的结果集,不需要保留全部原始数据(比如只需要学生ID和对应的等级,不需要原始分数)。
    • 减少数据框的列数,提升后续操作的效率(比如从宽表中仅提取关键衍生字段)。

内容的提问来源于stack exchange,提问作者Aswath Cm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 04:20:05