如何批量处理前缀为points_的考试得分变量转换为0/1对错标识?
points_* Variables to 0/1 (Correct/Incorrect) Got it, let's put an end to that tedious repetitive coding for your exam score variables! Since all your target variables follow the points_* naming pattern, we can use batch variable handling in whatever tool you’re using—no more writing code for each points_18616-style variable one by one. Here are tailored solutions for the most common data analysis tools:
R (Using Tidyverse/dplyr)
If you’re working in R with the tidyverse, dplyr’s across() function lets you target all points_ prefixed variables in one go:
library(dplyr) # Replace `df` with your actual dataset name df <- df %>% # Convert all points_* variables to 1 (if score > 0) or 0 (if score = 0) mutate( across( starts_with("points_"), ~ifelse(.x > 0, 1, 0), # Optional: Rename new variables to "correct_*" instead of overwriting originals .names = "correct_{col}" ) )
- If you want to overwrite the original
points_*variables instead of creating new ones, just remove the.namesargument. - Adjust the condition (
>.x > 0) if your "correct" definition differs (e.g.,.x == full_scoreif only perfect answers count).
Stata
In Stata, you can use a local macro to grab all matching variables, then loop through them:
// Step 1: Get a list of all variables starting with "points_" local point_vars: list varlist & points_* // Step 2: Loop through each variable to create 0/1 correct indicators foreach var of local point_vars { // Create a new variable (correct_points_*) with 1 for non-zero scores, 0 otherwise gen correct_`var' = (`var' > 0) // Uncomment below to overwrite the original `points_*` variable instead: // replace `var' = (`var' > 0) }
Stata automatically converts logical expressions (var > 0) to 1/0 values, so this is concise and efficient.
Python (Using Pandas)
For pandas users, filter the target columns and apply a vectorized transformation:
import pandas as pd # Replace `df` with your dataset name # Get all columns starting with "points_" point_columns = df.filter(like="points_").columns # Option 1: Create new "correct_*" columns df[[f"correct_{col}" for col in point_columns]] = df[point_columns].applymap(lambda x: 1 if x > 0 else 0) # Option 2: Overwrite the original `points_*` columns # df[point_columns] = df[point_columns].applymap(lambda x: 1 if x > 0 else 0)
applymap() applies the lambda function to every element in the selected columns, making this a fast, batch operation.
All these approaches work across different datasets as long as the points_* naming rule holds—no need to tweak code for each new dataset!
内容的提问来源于stack exchange,提问作者DrL

