You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何在for循环内调用DataFrame属性?自定义Bagging算法求助

Hey there! Since you're a Python newbie working on your thesis—needing a custom Bagging implementation (no off-the-shelf tools allowed) plus CSV splitting/merging—let's walk through this clearly, step by step.

Step 1: Reusable CSV Splitting & Merging (Beyond Your Specific File)

First, let's make your CSV handling code flexible so it works for any dataset with class labels. Let's assume your CSV has a column named category that holds your X class values.

import pandas as pd

# Load your input CSV
original_df = pd.read_csv("your_input_data.csv")

# Split into class-specific DataFrames (store in a dictionary for easy access later)
class_dataframes = {}
for class_label, group in original_df.groupby("category"):
    class_dataframes[class_label] = group
    # Optional: Save individual class CSVs if you need them for debugging
    group.to_csv(f"class_{class_label}.csv", index=False)

# Merge all class DataFrames back into a single CSV (order preserved by class label)
merged_df = pd.concat(class_dataframes.values(), ignore_index=True)
merged_df.to_csv("merged_classes_output.csv", index=False)

This code uses pandas groupby to split your data cleanly, stores each class's DataFrame in a dictionary (super useful for Bagging's per-class sampling later), and merges everything back seamlessly.

Step 2: Custom Bagging Implementation (No Sklearn Bagging Allowed)

Bagging's core is bootstrap sampling (random sampling with replacement) and aggregating predictions from multiple base models. Let's build this from scratch, using a decision tree as our base classifier (you can replace this with your own custom classifier if your thesis requires it).

First: Bootstrap Sampling Function

This function creates a random sample of your data (with replacement) — the foundation of Bagging:

import numpy as np

def bootstrap_sample(dataframe):
    # Generate random indices with replacement, matching the original data size
    sample_indices = np.random.choice(dataframe.index, size=len(dataframe), replace=True)
    return dataframe.loc[sample_indices]

Second: Full Custom Bagging Class

This class handles training multiple base models on bootstrap samples and making predictions via majority vote:

# If you can use a basic sklearn classifier as your base, use this; otherwise replace with your own!
from sklearn.tree import DecisionTreeClassifier

class CustomBaggingClassifier:
    def __init__(self, num_models=10, base_model=DecisionTreeClassifier()):
        self.num_models = num_models  # Number of base classifiers to train
        self.base_model = base_model
        self.trained_models = []  # Store all trained base models
    
    def fit(self, features, target):
        # Train num_models different base classifiers on bootstrap samples
        for _ in range(self.num_models):
            # For better performance (especially with imbalanced classes), use stratified sampling:
            # Sample from each class separately instead of the whole dataset
            bootstrap_features = []
            bootstrap_target = []
            unique_classes = np.unique(target)
            
            for cls in unique_classes:
                # Get indices of all samples in this class
                class_indices = np.where(target == cls)[0]
                # Bootstrap sample this class
                sampled_indices = np.random.choice(class_indices, size=len(class_indices), replace=True)
                bootstrap_features.extend(features[sampled_indices])
                bootstrap_target.extend(target[sampled_indices])
            
            # Convert back to numpy arrays for training
            bootstrap_features = np.array(bootstrap_features)
            bootstrap_target = np.array(bootstrap_target)
            
            # Train a new base model instance (critical: don't reuse the same model!)
            model_copy = self.base_model.__class__()
            model_copy.fit(bootstrap_features, bootstrap_target)
            self.trained_models.append(model_copy)
    
    def predict(self, features):
        # Get predictions from all trained models
        all_predictions = np.array([model.predict(features) for model in self.trained_models])
        # Majority vote for each input sample
        return np.array([np.bincount(preds).argmax() for preds in all_predictions.T])

Key Notes for Your Thesis:

  • If your thesis requires zero external classifier libraries, replace DecisionTreeClassifier with your own custom implementation (e.g., a hand-coded decision tree). This ensures your Bagging is fully original.
  • The stratified sampling in the fit method helps with imbalanced datasets — if your classes are evenly distributed, you can simplify to sampling the entire dataset at once instead of per-class.
Step 3: Tie It All Together with Your CSV Data

Let's use your CSV data to train and test the custom Bagging classifier:

# Load your data and split into features/target
df = pd.read_csv("your_input_data.csv")
# Adjust columns to match your dataset: drop non-feature columns (like category/ID)
features = df.drop(columns=["category", "target_label"]).values
target = df["target_label"].values  # Replace with your actual target column name

# Initialize and train your Bagging classifier
bagger = CustomBaggingClassifier(num_models=15)  # More models = more stable predictions
bagger.fit(features, target)

# Test with a sample from your data
test_sample = features[0].reshape(1, -1)  # Take the first row as a test example
predicted_class = bagger.predict(test_sample)
print(f"Predicted class: {predicted_class[0]}")
Bonus: Validate Your Implementation

For your thesis, you'll need to prove your Bagging works! Add evaluation code to test accuracy:

from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score

# Split data into train/test sets
X_train, X_test, y_train, y_test = train_test_split(features, target, test_size=0.2, random_state=42)

# Train on training data
bagger.fit(X_train, y_train)

# Predict on test data and calculate accuracy
y_pred = bagger.predict(X_test)
print(f"Bagging Classifier Accuracy: {accuracy_score(y_test, y_pred):.2f}")

内容的提问来源于stack exchange,提问作者Mateus Jose

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:50:11