如何为数据集新增列并填充预选随机值——TEAM数据集关联实践
Hey there! Let's tackle your two data manipulation questions with practical examples in both R and Python—two of the most popular tools for this kind of work.
1. General Method: Add a New Column with Predefined Random Values
The core idea here is to randomly select values from your predefined list for each row in your dataset. Here's how to do it in both languages:
Using R (with dplyr)
We'll use the sample() function to pick values from our list, and dplyr::mutate() to add the new column:
# Load the dplyr package (install first if needed: install.packages("dplyr")) library(dplyr) # Example dataset (replace with your actual data) your_dataset <- data.frame( ID = 1:8, Data_Point = rnorm(8) ) # Your predefined random values predefined_values <- c(10, 20, 30, 40) # Add the new random column your_dataset <- your_dataset %>% mutate(random_new_col = sample(predefined_values, nrow(your_dataset), replace = TRUE))
Note: Setting replace = TRUE lets us reuse values from the list if your dataset has more rows than the number of predefined values. If you don't want duplicates, set this to FALSE—just make sure your dataset has ≤ the number of values in your list!
Using Python (with pandas and numpy)
We'll use numpy.random.choice() to select values, then assign the result as a new column in our pandas DataFrame:
import pandas as pd import numpy as np # Example dataset (replace with your actual data) your_dataset = pd.DataFrame({ "ID": range(1, 9), "Data_Point": np.random.normal(size=8) }) # Your predefined random values predefined_values = [10, 20, 30, 40] # Add the new random column your_dataset["random_new_col"] = np.random.choice(predefined_values, size=len(your_dataset), replace=True)
Same note as above: replace=True allows duplicates, which is necessary if your dataset has more rows than your predefined value list.
2. Specific Task: Add Random League IDs to TEAM Dataset & Associate with LEAGUES
Let's walk through adding the random values from (1, 1729, 4679, 10257) to your TEAM dataset, then joining it with the LEAGUES dataset to link team and league info.
Using R
# Example TEAM dataset (replace with your actual data) TEAM <- data.frame( Team_ID = 1:12, Team_Name = paste0("Team_", 1:12) ) # Example LEAGUES dataset (replace with your actual data) LEAGUES <- data.frame( League_ID = c(1, 1729, 4679, 10257), League_Name = c("Top Flight", "Second Division", "Third Division", "Fourth Division") ) # Step 1: Add random League_ID column to TEAM TEAM <- TEAM %>% mutate(League_ID = sample(c(1, 1729, 4679, 10257), nrow(TEAM), replace = TRUE)) # Step 2: Join TEAM with LEAGUES to get league details TEAM_LEAGUE_ASSOCIATION <- TEAM %>% left_join(LEAGUES, by = "League_ID")
The left_join() ensures we keep all rows from the TEAM dataset, even if there's no matching League_ID (though in this case, all our random values exist in LEAGUES, so every row will get a League_Name).
Using Python
import pandas as pd import numpy as np # Example TEAM dataset (replace with your actual data) TEAM = pd.DataFrame({ "Team_ID": range(1, 13), "Team_Name": [f"Team_{i}" for i in range(1, 13)] }) # Example LEAGUES dataset (replace with your actual data) LEAGUES = pd.DataFrame({ "League_ID": [1, 1729, 4679, 10257], "League_Name": ["Top Flight", "Second Division", "Third Division", "Fourth Division"] }) # Step 1: Add random League_ID column to TEAM TEAM["League_ID"] = np.random.choice([1, 1729, 4679, 10257], size=len(TEAM), replace=True) # Step 2: Merge TEAM with LEAGUES to associate league info TEAM_LEAGUE_ASSOCIATION = pd.merge(TEAM, LEAGUES, on="League_ID", how="left")
The how="left" parameter here does the same thing as R's left_join()—preserves all rows from TEAM while matching league data where possible.
内容的提问来源于stack exchange,提问作者Justin Coleman

