Python中如何在for循环内调用DataFrame属性?自定义Bagging算法求助
Hey there! Since you're a Python newbie working on your thesis—needing a custom Bagging implementation (no off-the-shelf tools allowed) plus CSV splitting/merging—let's walk through this clearly, step by step.
First, let's make your CSV handling code flexible so it works for any dataset with class labels. Let's assume your CSV has a column named category that holds your X class values.
import pandas as pd # Load your input CSV original_df = pd.read_csv("your_input_data.csv") # Split into class-specific DataFrames (store in a dictionary for easy access later) class_dataframes = {} for class_label, group in original_df.groupby("category"): class_dataframes[class_label] = group # Optional: Save individual class CSVs if you need them for debugging group.to_csv(f"class_{class_label}.csv", index=False) # Merge all class DataFrames back into a single CSV (order preserved by class label) merged_df = pd.concat(class_dataframes.values(), ignore_index=True) merged_df.to_csv("merged_classes_output.csv", index=False)
This code uses pandas groupby to split your data cleanly, stores each class's DataFrame in a dictionary (super useful for Bagging's per-class sampling later), and merges everything back seamlessly.
Bagging's core is bootstrap sampling (random sampling with replacement) and aggregating predictions from multiple base models. Let's build this from scratch, using a decision tree as our base classifier (you can replace this with your own custom classifier if your thesis requires it).
First: Bootstrap Sampling Function
This function creates a random sample of your data (with replacement) — the foundation of Bagging:
import numpy as np def bootstrap_sample(dataframe): # Generate random indices with replacement, matching the original data size sample_indices = np.random.choice(dataframe.index, size=len(dataframe), replace=True) return dataframe.loc[sample_indices]
Second: Full Custom Bagging Class
This class handles training multiple base models on bootstrap samples and making predictions via majority vote:
# If you can use a basic sklearn classifier as your base, use this; otherwise replace with your own! from sklearn.tree import DecisionTreeClassifier class CustomBaggingClassifier: def __init__(self, num_models=10, base_model=DecisionTreeClassifier()): self.num_models = num_models # Number of base classifiers to train self.base_model = base_model self.trained_models = [] # Store all trained base models def fit(self, features, target): # Train num_models different base classifiers on bootstrap samples for _ in range(self.num_models): # For better performance (especially with imbalanced classes), use stratified sampling: # Sample from each class separately instead of the whole dataset bootstrap_features = [] bootstrap_target = [] unique_classes = np.unique(target) for cls in unique_classes: # Get indices of all samples in this class class_indices = np.where(target == cls)[0] # Bootstrap sample this class sampled_indices = np.random.choice(class_indices, size=len(class_indices), replace=True) bootstrap_features.extend(features[sampled_indices]) bootstrap_target.extend(target[sampled_indices]) # Convert back to numpy arrays for training bootstrap_features = np.array(bootstrap_features) bootstrap_target = np.array(bootstrap_target) # Train a new base model instance (critical: don't reuse the same model!) model_copy = self.base_model.__class__() model_copy.fit(bootstrap_features, bootstrap_target) self.trained_models.append(model_copy) def predict(self, features): # Get predictions from all trained models all_predictions = np.array([model.predict(features) for model in self.trained_models]) # Majority vote for each input sample return np.array([np.bincount(preds).argmax() for preds in all_predictions.T])
Key Notes for Your Thesis:
- If your thesis requires zero external classifier libraries, replace
DecisionTreeClassifierwith your own custom implementation (e.g., a hand-coded decision tree). This ensures your Bagging is fully original. - The stratified sampling in the
fitmethod helps with imbalanced datasets — if your classes are evenly distributed, you can simplify to sampling the entire dataset at once instead of per-class.
Let's use your CSV data to train and test the custom Bagging classifier:
# Load your data and split into features/target df = pd.read_csv("your_input_data.csv") # Adjust columns to match your dataset: drop non-feature columns (like category/ID) features = df.drop(columns=["category", "target_label"]).values target = df["target_label"].values # Replace with your actual target column name # Initialize and train your Bagging classifier bagger = CustomBaggingClassifier(num_models=15) # More models = more stable predictions bagger.fit(features, target) # Test with a sample from your data test_sample = features[0].reshape(1, -1) # Take the first row as a test example predicted_class = bagger.predict(test_sample) print(f"Predicted class: {predicted_class[0]}")
For your thesis, you'll need to prove your Bagging works! Add evaluation code to test accuracy:
from sklearn.model_selection import train_test_split from sklearn.metrics import accuracy_score # Split data into train/test sets X_train, X_test, y_train, y_test = train_test_split(features, target, test_size=0.2, random_state=42) # Train on training data bagger.fit(X_train, y_train) # Predict on test data and calculate accuracy y_pred = bagger.predict(X_test) print(f"Bagging Classifier Accuracy: {accuracy_score(y_test, y_pred):.2f}")
内容的提问来源于stack exchange,提问作者Mateus Jose

