在R语言中重塑调查数据以统计响应次数
Hey there! Let's work through reshaping your survey data and counting how often each option was selected for every question in R. I'll cover two common approaches—one using the tidyverse (great for readability and follow-up analysis) and another using base R (no extra packages needed).
First, let's start with a complete version of your data frame (I filled in the missing values for Q6; just replace this with your actual Q6 data if needed):
# Define your full data frame df <- structure( list( User = c("user1", "user2", "user3", "user4", "user5", "user6", "user7", "user8", "user9", "user10", "user11", "user12", "user13", "user14"), Q1 = c(0, 3, 5, 0, 6, 5, 1, 4, 6, 4, 5, 0, 0, 0), Q2 = c(0, 6, 4, 0, 4, 5, 0, 4, 6, 5, 5, 5, 0, 4), Q3 = c(0, 4, 5, 3, 4, 5, 0, 4, 4, 5, 5, 0, 0, 0), Q4 = c(5L, 6L, 6L, 7L, 6L, 6L, 6L, 4L, 3L, 4L, 6L, 5L, 3L, 6L), Q5 = c(7L, 5L, 6L, 7L, 5L, 5L, 7L, 4L, 4L, 6L, 6L, 6L, 6L, 6L), Q6 = c(6, 5, 7, 7, 7, 6, 6, 5, 4, 5, 6, 6, 5, 7) ), class = "data.frame", row.names = c(NA, -14L) )
Approach 1: Using the Tidyverse (dplyr + tidyr)
This method is intuitive, produces clean output, and makes it easy to extend for further analysis or visualization.
# Load the tidyverse package (install it first if you haven't: install.packages("tidyverse")) library(tidyverse) # Reshape data to long format and count responses response_counts <- df %>% # Convert wide data to long format: create "question" and "response" columns pivot_longer( cols = starts_with("Q"), # Target all columns starting with "Q" names_to = "question", # Name for the column holding question IDs values_to = "response" # Name for the column holding response values ) %>% # Count how many times each response appears per question count(question, response, name = "count") %>% # Optional: Sort results for readability arrange(question, response) # View the final counts print(response_counts)
What each step does:
pivot_longer: Flattens your wide data into a long format where each row represents one user's response to one question. This is essential for grouping and counting.count: Groups the data by question and response value, then calculates the number of times each combination occurs.arrange: Sorts the results so responses for each question are ordered numerically, making it easier to scan.
Approach 2: Using Base R
If you prefer not to load external packages, this base R method works just as well for quick counts:
# Use table() to count responses, then convert to a data frame base_counts <- as.data.frame( table( question = factor(unlist(df[, starts_with("Q")]), levels = paste0("Q", 1:6)), response = unlist(df[, starts_with("Q")]) ) ) # Rename the count column for clarity names(base_counts)[3] <- "count" # View the results print(base_counts)
Notes:
- The
factor()part ensures questions are ordered Q1 to Q6 in the output, rather than alphabetical order. - This method gives you the same core counts as the tidyverse approach, but the tidyverse version is easier to modify if you want to filter out certain responses (e.g., remove 0s if they represent skipped questions) or visualize the data.
内容的提问来源于stack exchange,提问作者Jason Matney

