You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R或Python中构建以整份数据框为样本的分类模型

Table-Level Classification for Dataframes (R & Python Solutions)

Great question! This is a classic table-level (or sequence-level) classification task where each "sample" is an entire dataframe (a time-series table) rather than individual rows. The core idea is to extract meaningful features from each dataframe, then train a classifier on these features to predict labels for new dataframes. Let's walk through practical implementations in both R and Python.

General Workflow

Before diving into code, here's the high-level process:

  • Organize your data: Collect all labeled dataframes into a structured format (e.g., list of (dataframe, label) pairs).
  • Feature extraction: Convert each dataframe into a fixed-length feature vector (using statistical summaries, time-series patterns, or deep learning embeddings).
  • Build training dataset: Combine all feature vectors with their corresponding labels into a standard tabular dataset.
  • Train a classifier: Use traditional ML models (like random forests) or deep learning models (like LSTMs) depending on your data size and complexity.
  • Predict on new dataframes: Extract the same features from new dataframes and use your trained model to predict labels.

Python Implementation

We'll use pandas for data handling, scikit-learn for ML models, and start with simple statistical features (easy to implement and interpret).

Step 1: Prepare Sample Data

First, let's simulate some labeled dataframes:

import pandas as pd
import numpy as np
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score

# Simulate labeled dataframes
def create_sample_df(label):
    np.random.seed(hash(label) % 100)
    time = np.arange(1,5)
    if label == "l1":
        measure1 = np.random.randint(10,20, size=4)
        measure2 = np.random.randint(1000,2000, size=4)
    elif label == "l2":
        measure1 = np.random.randint(20,30, size=4)
        measure2 = np.random.randint(500,1000, size=4)
    else: # l3
        measure1 = np.random.randint(0,10, size=4)
        measure2 = np.random.randint(2000,3000, size=4)
    return pd.DataFrame({"Time": time, "Measure1": measure1, "Measure2": measure2})

# Create 10 samples for each label
labeled_dfs = []
for label in ["l1", "l2", "l3"]:
    for _ in range(10):
        labeled_dfs.append( (create_sample_df(label), label) )

Step 2: Define Feature Extraction Function

We'll extract key statistics from each dataframe (you can add more features like trends, quantiles, etc.):

def extract_features(df):
    # Drop Time column if it's just an index (adjust if Time has meaning)
    numeric_df = df.drop("Time", axis=1)
    # Calculate summary statistics
    features = {
        "mean_measure1": numeric_df["Measure1"].mean(),
        "std_measure1": numeric_df["Measure1"].std(),
        "max_measure1": numeric_df["Measure1"].max(),
        "min_measure1": numeric_df["Measure1"].min(),
        "trend_measure1": np.polyfit(df["Time"], df["Measure1"], 1)[0], # slope of linear trend
        "mean_measure2": numeric_df["Measure2"].mean(),
        "std_measure2": numeric_df["Measure2"].std(),
        "max_measure2": numeric_df["Measure2"].max(),
        "min_measure2": numeric_df["Measure2"].min(),
        "trend_measure2": np.polyfit(df["Time"], df["Measure2"], 1)[0]
    }
    return pd.Series(features)

Step 3: Build Training Dataset

# Extract features for all samples
X = pd.DataFrame([extract_features(df) for df, _ in labeled_dfs])
y = [label for _, label in labeled_dfs]

# Split into train/test sets
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

Step 4: Train & Evaluate Model

# Train a Random Forest classifier
model = RandomForestClassifier(n_estimators=100, random_state=42)
model.fit(X_train, y_train)

# Evaluate
y_pred = model.predict(X_test)
print(f"Test Accuracy: {accuracy_score(y_test, y_pred):.2f}")

Step 5: Predict on New Dataframe

# Create a new dataframe to classify
new_df = create_sample_df("l2") # Replace with your actual new dataframe
new_features = extract_features(new_df)
predicted_label = model.predict(new_features.values.reshape(1, -1))
print(f"Predicted Label: {predicted_label[0]}")

For more complex patterns, you could use deep learning (e.g., feed the sequence directly into an LSTM or Transformer model using TensorFlow/PyTorch).


R Implementation

We'll use tidyverse for data manipulation, ranger for fast random forests, and caret for model evaluation.

Step 1: Prepare Sample Data

library(tidyverse)
library(ranger)
library(caret)

# Simulate labeled dataframes
create_sample_df <- function(label) {
  set.seed(hash(label) %% 100)
  time <- 1:4
  if (label == "l1") {
    measure1 <- sample(10:19, 4, replace = TRUE)
    measure2 <- sample(1000:1999, 4, replace = TRUE)
  } else if (label == "l2") {
    measure1 <- sample(20:29, 4, replace = TRUE)
    measure2 <- sample(500:999, 4, replace = TRUE)
  } else { # l3
    measure1 <- sample(0:9, 4, replace = TRUE)
    measure2 <- sample(2000:2999, 4, replace = TRUE)
  }
  tibble(Time = time, Measure1 = measure1, Measure2 = measure2)
}

# Create 10 samples for each label
labeled_dfs <- list()
for (label in c("l1", "l2", "l3")) {
  for (i in 1:10) {
    labeled_dfs[[length(labeled_dfs)+1]] <- list(df = create_sample_df(label), label = label)
  }
}

Step 2: Define Feature Extraction Function

extract_features <- function(df) {
  # Calculate summary statistics and trend slopes
  df %>%
    summarize(
      mean_measure1 = mean(Measure1),
      std_measure1 = sd(Measure1),
      max_measure1 = max(Measure1),
      min_measure1 = min(Measure1),
      trend_measure1 = lm(Measure1 ~ Time, data = df)$coefficients[["Time"]],
      mean_measure2 = mean(Measure2),
      std_measure2 = sd(Measure2),
      max_measure2 = max(Measure2),
      min_measure2 = min(Measure2),
      trend_measure2 = lm(Measure2 ~ Time, data = df)$coefficients[["Time"]]
    )
}

Step 3: Build Training Dataset

# Extract features for all samples
training_data <- map_dfr(labeled_dfs, function(item) {
  extract_features(item$df) %>%
    mutate(label = item$label)
})

# Split into train/test sets
set.seed(42)
train_idx <- createDataPartition(training_data$label, p = 0.8, list = FALSE)
X_train <- training_data[train_idx, !names(training_data) %in% "label"]
y_train <- training_data$label[train_idx]
X_test <- training_data[-train_idx, !names(training_data) %in% "label"]
y_test <- training_data$label[-train_idx]

Step 4: Train & Evaluate Model

# Train Random Forest model
model <- ranger(label ~ ., data = training_data[train_idx, ], num.trees = 100, seed = 42)

# Evaluate
y_pred <- predict(model, data = X_test)$predictions
confusionMatrix(y_pred, y_test)

Step 5: Predict on New Dataframe

# Create a new dataframe to classify
new_df <- create_sample_df("l3") # Replace with your actual new dataframe
new_features <- extract_features(new_df)
predicted_label <- predict(model, data = new_features)$predictions
cat("Predicted Label:", predicted_label, "\n")

内容的提问来源于stack exchange,提问作者Wolf_Cola

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:45:25