如何对月度股票收益DataFrame按月份为股票收益排名并生成新DataFrame
Hey there! Let's get your stock return ranking sorted out. First off, we need to handle those percentage strings (like "1%") by converting them to numeric values—otherwise the ranking won't work right. Then we can generate the rank DataFrame you need. Here are a few solid approaches in R:
Method 1: Base R with apply() (Quick & Simple)
This uses base R functions to process each row directly:
# Step 1: Convert percentage strings to numeric values # Assuming your first column is the month (e.g., "Jun 1927") df[-1] <- lapply(df[-1], function(col) { as.numeric(sub("%", "", col)) / 100 }) # Step 2: Generate rankings row-wise df_rank <- data.frame( Month = df[, 1], # Use rank(-row) to get higher returns as lower rank numbers t(apply(df[-1], 1, function(row) rank(-row, ties.method = "min"))) ) # Match column names to original DataFrame colnames(df_rank) <- colnames(df)
The rank(-row) trick reverses the order so higher returns get the top (smallest) ranks. The ties.method = "min" handles ties by assigning the smallest possible rank to matching values—you can swap this with "average", "max", or "first" if you need a different tie-breaking rule.
Method 2: Tidyverse Approach (For dplyr Fans)
If you prefer the tidyverse workflow, this uses dplyr and tidyr to reshape and rank:
library(dplyr) library(tidyr) df_rank <- df %>% # Reshape to long format for easier grouping pivot_longer(-1, names_to = "Stock", values_to = "Return") %>% # Convert percentages to numeric mutate(Return = as.numeric(sub("%", "", Return)) / 100) %>% # Group by month to rank within each period group_by(!!sym(colnames(df)[1])) %>% # Assign ranks (higher returns = lower rank number) mutate(Rank = rank(-Return, ties.method = "min")) %>% # Reshape back to wide format matching your original structure pivot_wider(names_from = "Stock", values_from = "Rank")
Method 3: Using Your Initial Empty Matrix Idea
If you want to stick with the empty DataFrame you started building, here's how to fill it in:
# Initialize empty DataFrame with same dimensions as df df_rank <- data.frame(matrix(NA, nrow = nrow(df), ncol = ncol(df))) colnames(df_rank) <- colnames(df) # Fill in the month column first df_rank[, 1] <- df[, 1] # Loop through each row to calculate ranks for (i in 1:nrow(df)) { # Convert current row's returns to numeric row_returns <- as.numeric(sub("%", "", df[i, -1])) / 100 # Calculate and assign ranks df_rank[i, -1] <- rank(-row_returns, ties.method = "min") }
All these methods will give you the df_rank structure you showed in your example—with each month's stocks ranked by their returns (higher returns get lower rank numbers).
内容的提问来源于stack exchange,提问作者Niccola Tartaglia

