如何基于其他DataFrame的信息对R语言DataFrame列进行减法运算?
Hey there! Let's work through this problem together—since you're new to R, I'll break it down clearly so you understand each step. The goal is to add new columns to your base dataframe using the calculation rules defined in the key dataframe, right?
First, let's recap the requirements to make sure we're on the same page:
- Only rows in
keywhereinclude = "yes"need to be processed - For each valid row, create a new column (named from the
namesfield) that equals the value of thecolscolumn minus thesubtractcolumn inbase
Base R Approach (No Extra Packages Needed)
This method uses core R functions, which is great if you don't want to install additional packages. A common pitfall with loops here is using $ to access columns dynamically—instead, use [[ which works with variable column names.
# Your original data base <- data.frame("A"=c("orange","apple","banana"), "B"=c(5,3,6), "C"=c(7,12,4), "D"=c(5,2,7), "E"=c(1,18,4)) key <- data.frame("cols"=c("A","B","C","D","E"), "include"=c("no","no","yes","no","yes"), "subtract"=c("na","A","B","C","D"), "names"=c("na","G","H","I","J")) # Step 1: Filter only the rules we need (include = "yes") key_filtered <- key[key$include == "yes", ] # Step 2: Loop through each valid rule to create new columns for (i in 1:nrow(key_filtered)) { # Extract values from the filtered rule row current_col <- key_filtered$cols[i] subtract_col <- key_filtered$subtract[i] new_col_name <- key_filtered$names[i] # Calculate and assign the new column using [[ for dynamic access base[[new_col_name]] <- base[[current_col]] - base[[subtract_col]] } # Check the result print(base)
Running this will give you exactly the output dataframe you expected!
Tidyverse Approach (dplyr + purrr)
If you're open to using the tidyverse set of packages (which makes data manipulation more readable once you get the hang of it), here's a concise way to do this:
First, install and load the packages if you haven't already:
install.packages(c("dplyr", "purrr")) library(dplyr) library(purrr)
Then run this code:
# Your original data (same as before) base <- data.frame("A"=c("orange","apple","banana"), "B"=c(5,3,6), "C"=c(7,12,4), "D"=c(5,2,7), "E"=c(1,18,4)) key <- data.frame("cols"=c("A","B","C","D","E"), "include"=c("no","no","yes","no","yes"), "subtract"=c("na","A","B","C","D"), "names"=c("na","G","H","I","J")) # Filter valid rules and create new columns key_filtered <- key %>% filter(include == "yes") walk(1:nrow(key_filtered), function(i) { base <<- base %>% mutate(!!key_filtered$names[i] := .data[[key_filtered$cols[i]]] - .data[[key_filtered$subtract[i]]]) }) # View the result print(base)
Quick Explanation of Tidyverse Syntax:
filter(include == "yes"): Keeps only the rules we needwalk(): Iterates over each rule row without returning a value (perfect for modifying the dataframe)!!key_filtered$names[i]: Uses the value fromnamesas the new column name (called "unquoting").data[[col_name]]: Safely references columns in the dataframe when using variable names
Verify the Result
To make sure everything works as expected, you can compare your updated base dataframe to the desired output:
output <- data.frame("A"=c("orange","apple","banana"), "B"=c(5,3,6), "C"=c(7,12,4), "D"=c(5,2,7), "E"=c(1,18,4), "H"=c(2,9,-2), "J"=c(-4,16,-3)) all.equal(base, output) # Should return TRUE if everything matches
内容的提问来源于stack exchange,提问作者MVincent

