如何在R中依据NA占比规则计算心理测试子分数?
Hey there! Let's work through this PedsQL scoring problem together—handling missing values properly for psychological tests can feel overwhelming at first, but we'll break it down into manageable steps that make sense for a new R user.
First, Let's Anchor to Your Core Rules
Just to recap so we're aligned:
- If more than 50% of items in a dimension are missing, the subscore for that dimension should be
NA. - If 50% or more items are completed, replace missing items with the mean of the completed items for that row, then calculate the overall mean for the dimension.
Option 1: Clean, Readable Solution with dplyr (Recommended)
If you're open to using the dplyr package, this approach is straightforward and easy to tweak for other dimensions later. Let's use the physical health dimension as an example:
library(dplyr) # Calculate pedsql_phys with your NA handling rules d <- d %>% rowwise() %>% mutate( # Count how many items are missing in the physical dimension phys_missing = sum(is.na(c_across(pedsql_1:pedsql_8))), # Calculate the proportion of missing items phys_missing_prop = phys_missing / 8, # Apply your scoring logic pedsql_phys = if (phys_missing_prop >= 0.5) { NA_real_ # Too many missing values: set subscore to NA } else { # Compute mean, automatically ignoring NAs (which acts like imputing with row mean) mean(c_across(pedsql_1:pedsql_8), na.rm = TRUE) } ) %>% ungroup() %>% # Remove temporary helper columns if you don't need them select(-phys_missing, -phys_missing_prop)
Option 2: Batch Processing with a Loop (For Multiple Dimensions)
If you have several PedsQL dimensions to score (like emotional, social, or school functioning), a loop will save you from repeating code. Here's how to set it up:
# Define each dimension's item columns and total item count first pedsql_dimensions <- list( phys = list(items = paste0("pedsql_", 1:8), total_items = 8), emo = list(items = paste0("pedsql_", 9:14), total_items = 6), social = list(items = paste0("pedsql_", 15:18), total_items = 4), school = list(items = paste0("pedsql_", 19:24), total_items = 6) ) # Loop through each dimension to calculate subscores for (dim_name in names(pedsql_dimensions)) { dim_info <- pedsql_dimensions[[dim_name]] items <- dim_info$items total <- dim_info$total_items # Calculate missing proportion for each row missing_prop <- rowSums(is.na(d[, items])) / total # Compute subscore using your rules d[[paste0("pedsql_", dim_name)]] <- apply(d[, items], 1, function(row) { if (sum(is.na(row)) / total >= 0.5) { NA_real_ } else { mean(row, na.rm = TRUE) } }) }
Why psych::scoreItems Might Have Failed
You mentioned trying scoreItems without success—chances are you just needed to adjust the missing value parameters to match your 50% rule. Here's how to make it work correctly:
library(psych) # Define the key for the physical dimension phys_key <- list(phys = paste0("pedsql_", 1:8)) # Score with your exact NA rules phys_scores <- scoreItems( keys = phys_key, data = d, missing = "mean", # Impute missing values with the row's completed item mean delete = 0.5, # Set subscore to NA if >50% of items are missing impute = TRUE # Enable the imputation feature ) # Add the calculated score back to your data frame d$pedsql_phys <- phys_scores$scores
The delete = 0.5 parameter is the missing piece here—it tells the function to exclude any row where more than half the items are missing, which aligns perfectly with your rule.
Quick Tip for Verification
Always spot-check a few rows manually to make sure the scoring works as expected. For example:
- Pick a row with 3 missing items in the 8-item physical dimension (3/8 = 37.5% < 50%) and confirm the mean is calculated correctly.
- Pick a row with 5 missing items (5/8 = 62.5% > 50%) and confirm the subscore is
NA.
内容的提问来源于stack exchange,提问作者Rbeginner_psych

