如何批量为DataFrame的cc列生成累积和?技术求助
Hey there! Let's fix this cumulative sum issue for your cc columns. The key here is that you need to compute a running total for each cc column, and since you have an unknown number of these columns later on, we need a scalable approach that doesn't require hardcoding each column name.
First, let's break down why your previous apply attempts didn't work:
- In your functions, you tried to modify the original data frame (
df$x) directly instead of returning the processed column. When usingapplyon a data frame, each column is passed to your function as a vector, and your function should return the modified vector—applywill then reassemble these vectors into a new data frame. - You also referenced
df$xinside the function, which doesn't correspond to the current column being processed byapply; you should use the input parameterxinstead.
Let's start with a reproducible example of your data:
df1 <- data.frame( rr_1 = c(100, 200, 300, 400, 0), rr_2 = c(0, 100, 300, 500, 0), cc_1 = c(1, 1, 1, 1, 0), cc_2 = c(0, 1, 1, 1, 0) )
Method 1: Using dplyr (recommended for tidy data workflows)
If you're using dplyr (version 1.0.0 or newer), across() is perfect for batch-processing columns. We can target all columns starting with "cc", apply cumsum() to each, and update them in place:
library(dplyr) df_processed <- df1 %>% mutate(across(starts_with("cc"), cumsum))
Let's check the result:
print(df_processed) # rr_1 rr_2 cc_1 cc_2 # 1 100 0 1 0 # 2 200 100 2 1 # 3 300 300 3 2 # 4 400 500 4 3 # 5 0 0 4 3
This matches exactly what you're looking for! The cumsum() function naturally handles the final 0 values by carrying forward the last cumulative total, which is exactly your desired behavior.
Method 2: Using base R apply
If you prefer base R, we can adjust your approach to work correctly. The key is to have the function return the processed column instead of modifying the original data frame:
# Extract cc columns cc_cols <- df1 %>% select(starts_with("cc")) # Apply cumsum to each column; apply returns a matrix, so convert back to data frame cc_processed <- as.data.frame(apply(cc_cols, 2, cumsum)) # Combine with the non-cc columns df_processed_base <- cbind(df1 %>% select(-starts_with("cc")), cc_processed)
This will give you the same result as the dplyr method.
Why this works
Your target cumulative sequence is exactly what cumsum() produces:
- For
cc_1, the original values are[1,1,1,1,0]→cumsum()gives[1,2,3,4,4] - For
cc_2, original values are[0,1,1,1,0]→cumsum()gives[0,1,2,3,3]
No need for custom loops—R's built-in cumsum() does exactly what you need, and using across() or apply lets you apply it to all cc columns automatically, regardless of how many there are.
内容的提问来源于stack exchange,提问作者Luis Manso

