You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于其他DataFrame的信息对R语言DataFrame列进行减法运算?

Solution to Generate New Columns in R DataFrame Based on Rules

Hey there! Let's work through this problem together—since you're new to R, I'll break it down clearly so you understand each step. The goal is to add new columns to your base dataframe using the calculation rules defined in the key dataframe, right?

First, let's recap the requirements to make sure we're on the same page:

  • Only rows in key where include = "yes" need to be processed
  • For each valid row, create a new column (named from the names field) that equals the value of the cols column minus the subtract column in base

Base R Approach (No Extra Packages Needed)

This method uses core R functions, which is great if you don't want to install additional packages. A common pitfall with loops here is using $ to access columns dynamically—instead, use [[ which works with variable column names.

# Your original data
base <- data.frame("A"=c("orange","apple","banana"), "B"=c(5,3,6), "C"=c(7,12,4), "D"=c(5,2,7), "E"=c(1,18,4))
key <- data.frame("cols"=c("A","B","C","D","E"), "include"=c("no","no","yes","no","yes"), "subtract"=c("na","A","B","C","D"), "names"=c("na","G","H","I","J"))

# Step 1: Filter only the rules we need (include = "yes")
key_filtered <- key[key$include == "yes", ]

# Step 2: Loop through each valid rule to create new columns
for (i in 1:nrow(key_filtered)) {
  # Extract values from the filtered rule row
  current_col <- key_filtered$cols[i]
  subtract_col <- key_filtered$subtract[i]
  new_col_name <- key_filtered$names[i]
  
  # Calculate and assign the new column using [[ for dynamic access
  base[[new_col_name]] <- base[[current_col]] - base[[subtract_col]]
}

# Check the result
print(base)

Running this will give you exactly the output dataframe you expected!


Tidyverse Approach (dplyr + purrr)

If you're open to using the tidyverse set of packages (which makes data manipulation more readable once you get the hang of it), here's a concise way to do this:

First, install and load the packages if you haven't already:

install.packages(c("dplyr", "purrr"))
library(dplyr)
library(purrr)

Then run this code:

# Your original data (same as before)
base <- data.frame("A"=c("orange","apple","banana"), "B"=c(5,3,6), "C"=c(7,12,4), "D"=c(5,2,7), "E"=c(1,18,4))
key <- data.frame("cols"=c("A","B","C","D","E"), "include"=c("no","no","yes","no","yes"), "subtract"=c("na","A","B","C","D"), "names"=c("na","G","H","I","J"))

# Filter valid rules and create new columns
key_filtered <- key %>% filter(include == "yes")

walk(1:nrow(key_filtered), function(i) {
  base <<- base %>%
    mutate(!!key_filtered$names[i] := .data[[key_filtered$cols[i]]] - .data[[key_filtered$subtract[i]]])
})

# View the result
print(base)

Quick Explanation of Tidyverse Syntax:

  • filter(include == "yes"): Keeps only the rules we need
  • walk(): Iterates over each rule row without returning a value (perfect for modifying the dataframe)
  • !!key_filtered$names[i]: Uses the value from names as the new column name (called "unquoting")
  • .data[[col_name]]: Safely references columns in the dataframe when using variable names

Verify the Result

To make sure everything works as expected, you can compare your updated base dataframe to the desired output:

output <- data.frame("A"=c("orange","apple","banana"), "B"=c(5,3,6), "C"=c(7,12,4), "D"=c(5,2,7), "E"=c(1,18,4), "H"=c(2,9,-2), "J"=c(-4,16,-3))
all.equal(base, output) # Should return TRUE if everything matches

内容的提问来源于stack exchange,提问作者MVincent

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 08:07:43