如何在Python中对CSV数据集进行80-20训练测试拆分
Hey there! Let's sort out that dataset splitting problem you're working on. The code snippet you shared has a key issue—using open() and split() on the CSV file doesn't split the rows of your dataset (it just splits raw text into chunks), so that approach won't give you the proper 80/20 split you need.
Since you're already using pandas to load the data, here are two reliable, straightforward methods to split your dataset correctly:
Method 1: Using Pandas Built-in Sampling
If you want to keep the full DataFrame structure (features + target) for both sets, this works great:
import pandas as pd # Load your dataset (you already have this part right!) diabetes_df = pd.read_csv("diabetes.csv") # Set a random seed to make your split reproducible (so you get the same split every time) random_seed = 42 # Sample 80% of the data for training train_df = diabetes_df.sample(frac=0.8, random_state=random_seed) # The remaining 20% becomes the test set (drop the rows already in train) test_df = diabetes_df.drop(train_df.index) # Quick check to confirm the split sizes print(f"Training set: {len(train_df)} rows ({len(train_df)/len(diabetes_df)*100:.1f}%)") print(f"Test set: {len(test_df)} rows ({len(test_df)/len(diabetes_df)*100:.1f}%)")
Method 2: Using Scikit-learn's train_test_split (Recommended for Modeling)
Since you're planning to implement Logistic Regression, Naive Bayes, and Decision Trees, this method is better—it splits your data into separate feature matrices (X) and target vectors (y), which is exactly what scikit-learn's model APIs expect:
import pandas as pd from sklearn.model_selection import train_test_split diabetes_df = pd.read_csv("diabetes.csv") # Separate features (all columns except the target) and target variable # Replace 'Outcome' with your actual target column name (common in diabetes datasets) X = diabetes_df.drop('Outcome', axis=1) y = diabetes_df['Outcome'] # Split into train and test sets: 20% of data goes to test # `stratify=y` ensures the class distribution in train/test matches the original dataset # (super important if your target is imbalanced, e.g., more non-diabetic than diabetic samples) X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y ) # Verify the splits print(f"Training features shape: {X_train.shape}") print(f"Test features shape: {X_test.shape}") print(f"Training target shape: {y_train.shape}") print(f"Test target shape: {y_test.shape}")
Why stratify=y Matters?
If your target variable (like diabetes diagnosis) is imbalanced (e.g., only 30% of samples are positive), using stratify makes sure both your training and test sets have the same proportion of positive/negative samples. This prevents your model from being biased towards the majority class during training.
Either method will work, but the scikit-learn approach will save you extra steps when you start training your models later.
内容的提问来源于stack exchange,提问作者Georges Ridgmont

