R语言中使用dplyr对含NA值的列做除法失败,求解决方法
Hey there! Let's sort out that division problem you're having with NA values in your data frame. The trouble with your original code df %>% mutate(c_n = c / n, na.rm = TRUE) is that na.rm = TRUE isn't a valid argument for mutate() itself — that parameter works with aggregation functions like sum() or mean(), but not directly with basic arithmetic operations like division.
Here are a few practical solutions to handle NA values correctly during the division:
Solution 1: Skip Rows with NA Using ifelse()
This approach checks if either c or n is NA before performing the division. If either value is missing, it returns NA for the result; otherwise, it calculates the division:
library(dplyr) df %>% mutate(c_n = ifelse(is.na(c) | is.na(n), NA, c / n))
Solution 2: Clean NA Values First (Use Carefully)
If you want to replace NA values with a default value before dividing (make sure this makes sense for your data!), you can use replace_na(). For example, if you want to treat missing c as 0 and avoid division by zero by replacing missing n with 1:
df %>% mutate( c_clean = replace_na(c, 0), n_clean = replace_na(n, 1), # Adjust this default based on your data logic c_n = c_clean / n_clean ) %>% select(-c_clean, -n_clean) # Optional: Remove temporary columns
Solution 3: Row-Wise Check with rowwise()
If you prefer a more explicit row-wise check, you can use rowwise() to evaluate each row individually:
df %>% rowwise() %>% mutate(c_n = if(sum(is.na(c), is.na(n)) == 0) c/n else NA) %>% ungroup()
All these methods will correctly handle NA values in your c and n columns, so you'll get valid division results where both values exist, and NA where either is missing.
内容的提问来源于stack exchange,提问作者Tiptop

