You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言中Wilcoxon秩和检验的相同值修正秩次计算方法

Manual Wilcoxon Rank Sum Test: Handling Tied Values with Average Ranks

Hey there! I see you're working on manually implementing the Wilcoxon Rank Sum Test and need to fix the ranks for tied values (assigning the average rank to all rows with the same total value). Let's work through this step by step—first, I'll note your code was missing the definition for x, so I'll assume it's x <- 1:10 based on your output dataframe.

Solution 1: Using Base R's Built-in rank() Function (Simpler)

Even though you mentioned you know about built-in functions, this is the easiest way to get the average ranks you need, and it aligns cleanly with your workflow:

# Define your vectors first
x <- 1:10
y <- c(3,4,4,4,5,6,7)

# Create the initial dataframe
df <- data.frame(
  names = c(rep("X", length(x)), rep("Y", length(y))),
  total = c(x, y)
)

# Sort by total and calculate average ranks
rankdf <- df[order(df$total), ]
rankdf$rank <- rank(rankdf$total, ties.method = "average")

Solution 2: Fully Manual Calculation (Matching Your Original Approach)

If you want to avoid relying on the rank() function's built-in tie handling, here's how to compute average ranks manually:

# Start with your sorted dataframe
sorted_df <- df[order(df$total), ]
# Add a temporary sequential rank column (like your original `rank` column)
sorted_df$temp_rank <- 1:nrow(sorted_df)

# Calculate average ranks for each tied group using base R
rank_list <- tapply(sorted_df$temp_rank, sorted_df$total, mean)
sorted_df$rank <- rank_list[as.character(sorted_df$total)]

# Clean up the temporary column
rankdf <- sorted_df[, c("names", "total", "rank")]

Bonus: Readable Dplyr Alternative

If you're open to using the dplyr package (great for data manipulation beginners), this version is more intuitive:

library(dplyr)

rankdf <- df %>%
  arrange(total) %>%
  group_by(total) %>%
  mutate(rank = mean(1:nrow(df))[match(total, unique(total))]) %>%
  ungroup()

Final Result

All methods will produce exactly the dataframe you're expecting:

names total rank
1     X     1  1.0
2     X     2  2.0
3     X     3  3.5
11    Y     3  3.5
4     X     4  6.0
12    Y     4  6.0
13    Y     4  6.0
5     X     5  8.5
14    Y     5  8.5
6     X     6 10.5
15    Y     6 10.5
7     X     7 12.5
16    Y     7 12.5
8     X     8 14.0
9     X     9 15.0
10    X    10 16.0

How It Works

For tied values (like total=3), we take all sequential ranks assigned to those rows (3 and 4), compute their mean (3.5), and assign that mean to every row in the tied group. This follows Wilcoxon Rank Sum Test rules for handling ties, which is critical for calculating an accurate test statistic later.

内容的提问来源于stack exchange,提问作者Valentina Grillo-Alvarado

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 14:42:56