You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

生成特定性能数据集求助:适配LogisticRegression与DecisionTreeClassifier

Solution: Generate Datasets for Targeted Classifier Performance

Let's tackle this problem by designing each dataset to exploit the strengths and weaknesses of Logistic Regression and a depth-3 Decision Tree. I'll walk you through the logic behind each dataset and provide full Python code to generate, train, and evaluate everything.

Dataset 1: Logistic Regression Shines, Decision Tree Struggles

Core Idea

Logistic Regression excels at capturing global linear trends, while a shallow Decision Tree (depth=3) gets tripped up by scattered local noise. We'll build a linearly separable base dataset, then add spread-out noise points that the tree overfits to, but Logistic Regression ignores.

Code Implementation

import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score

# Generate Dataset 1 (reproducible results)
np.random.seed(42)
# Linearly separable base data: class 1 where x0 + x1 > 0
X_base = np.random.randn(1000, 2)
y_base = (X_base[:, 0] + X_base[:, 1] > 0).astype(int)
# Add scattered noise: flip labels for 30% of random samples
noise_indices = np.random.choice(1000, 300, replace=False)
y_noisy = y_base.copy()
y_noisy[noise_indices] = 1 - y_noisy[noise_indices]
X1, y1 = X_base, y_noisy

# Train and evaluate models
lr1 = LogisticRegression()
dt1 = DecisionTreeClassifier(max_depth=3)
lr1.fit(X1, y1)
dt1.fit(X1, y1)

print(f"Dataset 1 - Logistic Regression Accuracy: {accuracy_score(y1, lr1.predict(X1)):.2f}")
print(f"Dataset 1 - Decision Tree Accuracy: {accuracy_score(y1, dt1.predict(X1)):.2f}")

Why It Works

  • Logistic Regression focuses on the underlying linear boundary (x0 + x1 > 0) and is robust to scattered noise, so accuracy stays above 0.9.
  • The shallow Decision Tree can't tell the global trend apart from noise—it wastes splits on noisy points, leading to accuracy below 0.7.

Dataset 2: Decision Tree Shines, Logistic Regression Struggles

Core Idea

Decision Trees (even shallow ones) handle axis-aligned nonlinear patterns well, while Logistic Regression can only model linear boundaries. We'll create a dataset where classes live in diagonal quadrants—impossible for a linear model to split, but trivial for a depth-3 tree.

Code Implementation

# Generate Dataset 2
np.random.seed(42)
# Quadrant-based data: class 1 where x0 and x1 have the same sign
X2 = np.random.randn(1000, 2)
y2 = ((X2[:, 0] * X2[:, 1]) > 0).astype(int)

# Train and evaluate models
lr2 = LogisticRegression()
dt2 = DecisionTreeClassifier(max_depth=3)
lr2.fit(X2, y2)
dt2.fit(X2, y2)

print(f"\nDataset 2 - Logistic Regression Accuracy: {accuracy_score(y2, lr2.predict(X2)):.2f}")
print(f"Dataset 2 - Decision Tree Accuracy: {accuracy_score(y2, dt2.predict(X2)):.2f}")

Why It Works

  • Logistic Regression can only learn a straight line boundary, which can't separate diagonal quadrants—accuracy hovers around 0.5, well below 0.7.
  • The depth-3 Decision Tree splits on x0 and x1 thresholds, easily capturing the quadrant pattern, leading to accuracy above 0.9.

Dataset 3: Both Classifiers Struggle

Core Idea

We need a dataset with no meaningful pattern between features and labels. By creating completely mixed data with random labels, neither model can learn a useful boundary.

Code Implementation

# Generate Dataset 3
np.random.seed(42)
# Random feature data with no structure
X3 = np.random.randn(1000, 2)
# Random labels (500 of each class, shuffled)
y3 = np.concatenate([np.zeros(500), np.ones(500)])
np.random.shuffle(y3)

# Train and evaluate models
lr3 = LogisticRegression()
dt3 = DecisionTreeClassifier(max_depth=3)
lr3.fit(X3, y3)
dt3.fit(X3, y3)

print(f"\nDataset 3 - Logistic Regression Accuracy: {accuracy_score(y3, lr3.predict(X3)):.2f}")
print(f"Dataset 3 - Decision Tree Accuracy: {accuracy_score(y3, dt3.predict(X3)):.2f}")

Why It Works

  • There's no relationship between features and labels—both models can't learn anything meaningful. Their accuracy will be close to 0.5, well below 0.7.

内容的提问来源于stack exchange,提问作者Alex Nikitin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:59:46