如何用dplyr计算多行平均值?结果不符问题求助
Hey there! Let's dig into why your dplyr calculation is spitting out 526 instead of the 207.404 you calculated manually—this kind of big gap almost always comes down to a small, easy-to-fix misstep. Here are the most common culprits to check:
1. You used sum() instead of mean()
This is the #1 mistake I see for this exact issue! If your code looks something like this:
df %>% filter(your_row_filter_here) %>% summarise(result = sum(your_target_column))
You’re getting the total sum of the values instead of their average. Swap sum() for mean()—and don’t forget to add na.rm = TRUE if there are any missing values, since those can throw off results silently.
2. Incorrect filtering or accidental grouping
Double-check that your filter() statement is selecting the exact rows you used for your manual calculation. Maybe you included extra rows you didn’t intend, or missed some key ones? For example, if you filtered for region == "North" but your manual math used "South" rows, that’d skew the average hard.
Also, watch out for unintended group_by() calls! If you grouped by a column that splits your data into multiple groups, dplyr might be calculating averages per group instead of the overall average for your selected rows. If you don’t need grouping, either remove the group_by() line or add ungroup() before summarizing.
3. You picked the wrong column
It sounds silly, but it happens all the time! Confirm that the column name in your mean() call matches the one you used for your manual calculation. Maybe you referenced total_sales instead of unit_price, or a column with completely unrelated values.
4. Data type quirks (less likely, but worth verifying)
If your target column is stored as a character type instead of numeric, mean() would normally return NA—but if there was a weird coercion issue (like numbers stored as strings with extra characters), you might get unexpected results. Run class(df$your_target_column) to check if it’s numeric. If not, convert it first:
df %>% mutate(your_target_column = as.numeric(your_target_column)) %>% filter(your_row_filter_here) %>% summarise(average = mean(your_target_column, na.rm = TRUE))
Example of correct code
If you’re trying to calculate the average of rows 1-3 in the value column, it should look like this:
library(dplyr) your_data %>% filter(row_number() %in% 1:3) # Or use your specific row identifiers summarise(manual_match_average = mean(value, na.rm = TRUE))
Go through these checks one by one—chances are you’ll spot the issue in no time!
内容的提问来源于stack exchange,提问作者nik

