如何在R中打乱数据集的Day列值并实现批量多次打乱及结果存储?
Got it, let's solve this. Since your dataset has uneven observations per day, we can't just shuffle the Day column row-by-row—we need to reassign entire day groups to new labels. Here's a straightforward way to do this with Python and pandas, which will let you generate multiple shuffled datasets easily:
The core idea is to create random permutations of your unique day values (1 to 10), then map each original day to a new shuffled day as a group. This keeps all rows from the same original day together under a new day label, without altering any of the behavior data.
Step 1: Set Up Dependencies
First, make sure you have pandas and numpy installed (they’re standard for data tasks):
import pandas as pd import numpy as np import os
Step 2: Define the Shuffling Function
This function will handle loading your data, generating shuffled day mappings, and saving each result as a separate file:
def shuffle_day_groups(input_file_path, num_shuffles, output_folder): # Create output folder if it doesn't exist os.makedirs(output_folder, exist_ok=True) # Load your original dataset (adjust read method if using Excel/Parquet/etc.) original_data = pd.read_csv(input_file_path) # Get sorted unique day values (ensures we work with 1-10 consistently) unique_days = sorted(original_data['Day'].unique()) for shuffle_num in range(1, num_shuffles + 1): # Generate a random permutation of the unique days (this is the shuffle!) shuffled_day_order = np.random.permutation(unique_days) # Create a mapping: original day → new shuffled day day_mapping = dict(zip(unique_days, shuffled_day_order)) # Apply the mapping to copy of the original data (so we don't overwrite it) shuffled_data = original_data.copy() shuffled_data['Day'] = shuffled_data['Day'].map(day_mapping) # Save the shuffled dataset with a unique name output_file = os.path.join(output_folder, f"shuffled_data_{shuffle_num}.csv") shuffled_data.to_csv(output_file, index=False) print(f"Shuffle {shuffle_num} saved to: {output_file}")
Step 3: Run the Function
Call the function with your specific details:
# Example usage (replace with your file paths and desired number of shuffles) shuffle_day_groups( input_file_path="your_original_dataset.csv", num_shuffles=5, # How many shuffled datasets you want to create output_folder="shuffled_datasets" )
Key Details:
- Group Consistency: All rows originally labeled "Day 2" will get the same new day value in each shuffle (e.g., all become "Day 10" in one run, "Day 3" in another).
- Reproducibility: If you want consistent shuffles for testing, add
np.random.seed(42)before generating the permutation (inside the loop for unique seeds per shuffle, or outside for the same seed across all shuffles). - Flexibility: If your data isn’t in CSV, swap
pd.read_csv()andto_csv()withpd.read_excel()/to_excel()or other pandas read/write methods. - Non-Numeric Days: If your days are strings (like "Day1"), the code still works—pandas will handle string mappings the same way.
This approach keeps your behavior data intact while shuffling day labels exactly as you need, even with an imbalanced dataset.
内容的提问来源于stack exchange,提问作者Romulo

