如何基于现有数据集的变量与行创建新数据集?
Hey there! Let's break down how to create your new dataset using the variables you've already calculated, plus existing columns from your esports data. I'll cover a couple of common approaches depending on your needs:
1. Basic R Approach (No Extra Packages Needed)
If you prefer sticking to base R, these steps will work smoothly:
- Combine selected original columns with your new variables:
First pick the columns you want to keep from the originalesportsdataset (e.g.,Age,gender), then merge them withfemininityandmac:# Select specific columns from the original dataset original_subset <- esports[, c("Age", "gender")] # Create the new dataset by combining the subset with your calculated variables new_dataset <- data.frame(original_subset, femininity, mac) - Add new variables to the full original dataset:
If you want to keep all existing columns plus your new variables, usecbind():new_dataset_full <- cbind(esports, femininity, mac) - Filter rows based on conditions:
To include only specific rows (e.g., players over 18), subset the dataset like this:filtered_dataset <- new_dataset[new_dataset$Age > 18, ]
2. Tidyverse/dplyr Approach (More Intuitive for Data Manipulation)
If you're open to using the dplyr package (part of the tidyverse), this method is cleaner for complex data tasks:
First, make sure you have the package installed and loaded:
install.packages("dplyr") library(dplyr)
- Option 1: Calculate variables and build the dataset in one step:
You can even embed yourrowMeanscalculations directly into the workflow to avoid creating separate vectors first:new_dataset <- esports %>% # Add your calculated variables as new columns mutate( femininity = rowMeans(data.frame(.[11:15], 6 - .[16:17])), mac = rowMeans(data.frame(.[28:30], 6 - .[, c(31, 33)])) # Note: Fixed your original code here - use .[, c(31,33)] to get all rows for columns 31 and 33 ) %>% # Select only the columns you need (remove this line if you want all columns) select(Age, gender, femininity, mac) - Option 2: Bind pre-calculated variables to a subset of original data:
If you already havefemininityandmacready, usebind_cols()to merge them with your chosen original columns:new_dataset <- esports %>% select(Age, gender) %>% # Pick columns from original data bind_cols(femininity = femininity, mac = mac) - Filter rows with
filter():
Add afilter()step to keep only rows that meet your criteria:filtered_dataset <- new_dataset %>% filter(Age > 18, gender == "Female") # Example conditions - adjust as needed
Quick Note on Your Original mac Calculation
I noticed a small syntax issue in your original code: 6-esports[31,33] will only select the value at row 31, column 33. To get all rows for columns 31 and 33, you need to write 6-esports[, c(31, 33)] (the comma tells R to include all rows for those columns). This fix ensures your mac variable calculates correctly across all observations.
内容的提问来源于stack exchange,提问作者Fares Hassan Almas

