在R中针对实验多选择题的字符串型作答数据实现自动评分的方法
Here's a practical, step-by-step implementation to calculate category-based total scores using keyword matching (like option letters) for your experiment data:
Step 1: Load and Prepare Data
First, we’ll read your CSV file and separate the answer key (first row) from participant responses.
# Replace with your actual file path experiment_data <- read.csv("your_experiment_data.csv", stringsAsFactors = FALSE) # Extract the answer key (first row holds correct answers) answer_key <- experiment_data[1, ] # Isolate participant responses (all rows after the first) participant_responses <- experiment_data[-1, ]
Step 2: Define Category Prefixes
List the category prefixes you need to score:
category_prefixes <- c("PRE_TR", "PRE_IS", "PRE_RULE", "POST_TR", "POST_IS", "POST_RULE")
Step 3: Create a Scoring Helper Function
This function calculates total correct responses for a single category and participant. It extracts the option letter (e.g., A, B, C) from the correct answer and checks if the participant’s response contains that keyword.
# Load stringr for robust keyword extraction (optional but recommended) library(stringr) calculate_category_score <- function(participant_row, answer_key, category_prefix) { # Filter columns belonging to the current category category_cols <- grepl(paste0("^", category_prefix), names(participant_row)) # Get participant's answers and correct answers for these columns participant_answers <- participant_row[category_cols] correct_answers <- answer_key[category_cols] # Extract the option letter (works for formats like "A. Correct Option" or "(B) Another Option") correct_keywords <- str_extract(correct_answers, "[A-Z]") # Check if each response contains the correct keyword # Set ignore.case = TRUE if responses might use lowercase letters is_correct <- mapply(function(resp, key) { !is.na(resp) && grepl(key, resp, ignore.case = FALSE) }, participant_answers, correct_keywords) # Sum the number of correct responses sum(is_correct) }
Step 4: Calculate Scores for All Participants and Categories
Apply the helper function to generate total scores for each category:
# Calculate category totals for all participants category_scores <- as.data.frame(sapply(category_prefixes, function(prefix) { apply(participant_responses, 1, calculate_category_score, answer_key = answer_key, category_prefix = prefix) })) # Rename columns to match your desired output (e.g., TOTAL_PRE_TR) colnames(category_scores) <- paste0("TOTAL_", colnames(category_scores))
Step 5: Combine Data and Export
Merge the original participant data with the new total scores and save the result:
# Combine responses and scores final_scored_data <- cbind(participant_responses, category_scores) # Export to CSV (optional) write.csv(final_scored_data, "scored_experiment_data.csv", row.names = FALSE)
Key Notes & Adjustments
- Keyword Extraction: If your correct answers use a different format (e.g., lowercase letters, numbers), modify the
str_extractpattern (e.g.,[a-z]for lowercase,[0-9]for numbers). - NA Handling: The function treats unanswered questions (NA) as incorrect. If you want to exclude them from scoring, add
na.rm = TRUEto thesum()call. - Case Sensitivity: Change
ignore.case = FALSEtoTRUEif participant responses might use lowercase letters for options.
Example Output Structure
Your final data will include all original participant columns plus new columns like:
| ... | TOTAL_PRE_TR | TOTAL_PRE_IS | TOTAL_PRE_RULE | TOTAL_POST_TR | ... |
|---|---|---|---|---|---|
| 5 | 3 | 4 | 6 | ||
| 4 | 2 | 5 | 7 |
内容的提问来源于stack exchange,提问作者gokgok

