You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DNN准确率过高问题排查、训练错误定位及跨数据集测试咨询

Hey there! Let's tackle your two issues one by one—first that suspiciously perfect accuracy (which is definitely a red flag, not a win), then how to test your model on other datasets.

1. Diagnosing the "Too Good to Be True" Accuracy

Let's start with the root cause of your perfect metrics:

a. Critical Label Encoding Bug

Look at your label_fix function:

def label_fix(label):
    if label == '<=50K':
        return 0
    else:
        return 1

The US Adult Income dataset uses raw <=50K and >50K as income labels, not the HTML-escaped &lt;=50K. That means every single sample in your dataset is being labeled as 1—there are no 0s at all!

When your model trains on a dataset where all samples belong to one class, it can trivially predict that class every time and get 100% accuracy. That's why your loss drops to 0 instantly and your test metrics are perfect.

b. You're Using a Linear Model, Not a DNN

You mentioned training a DNN, but your code uses tf.estimator.LinearClassifier—this is a linear classification model, not a deep neural network. If you want a DNN, you'll need to switch to tf.estimator.DNNClassifier (and transform categorical columns first, since DNNs can't accept raw categorical inputs).

Fixes to Apply

  • Fix the label encoding: Update your function to match the actual labels in the dataset:
    def label_fix(label):
        if label == '<=50K':
            return 0
        else:
            return 1
    
    Or use pandas' cleaner map function:
    input_data['Income'] = input_data['Income'].map({'<=50K': 0, '>50K': 1})
    
  • Verify label distribution: After fixing the encoding, double-check that you have both classes:
    print(input_data['Income'].value_counts())
    
    You should see roughly 75% of samples labeled 0 and 25% labeled 1 (the standard split for this dataset).
  • Switch to a DNN (if desired): To use a DNNClassifier, convert categorical columns to embedding or indicator columns first:
    # Convert high-cardinality categorical columns to embeddings
    job_class_emb = tf.feature_column.embedding_column(Job_class, dimension=8)
    education_emb = tf.feature_column.embedding_column(Education, dimension=8)
    # Convert low-cardinality columns to indicator columns
    gender_ind = tf.feature_column.indicator_column(Gender)
    colour_ind = tf.feature_column.indicator_column(Colour)
    
    # Update your feature columns list
    feats_cols = [Age, fnlwgt, Education_num, Capital_gain, Capital_loss, Hours,
                  job_class_emb, education_emb, gender_ind, colour_ind, Status, Designation, Marital, Native_country]
    
    # Initialize the DNNClassifier
    model = tf.estimator.DNNClassifier(
        feature_columns=feats_cols,
        hidden_units=[256, 128, 64],  # Adjust these sizes based on your needs
        n_classes=2
    )
    
2. Testing Your Model on Another Dataset

Once you have a working model, testing it on new data requires aligning the new dataset with your training data's structure and preprocessing. Here's how:

Step 1: Match Feature Structure

Your new dataset must have all the same feature columns as your training data (exact names, matching data types). For example, if your training data uses Age and Job-Class, the new data can't have age or Job Class—stick to the exact column names.

Step 2: Apply Identical Preprocessing

  • Categorical features: Ensure the new data uses the same vocabulary as your training data. If the new data has unseen categories (e.g., a gender not present in training), map them to an "Unknown" category or remove those samples.
  • Label encoding: If you're evaluating on labeled new data, apply the same label mapping you used for the training set.

Step 3: Create an Input Function for the New Data

Use the same pandas_input_fn pattern as your test set:

# Load and preprocess new data
new_data = pd.read_csv('new_dataset.csv')
new_data['Income'] = new_data['Income'].map({'<=50K': 0, '>50K': 1})  # Skip if no labels
x_new = new_data.drop('Income', axis=1)
y_new = new_data['Income']  # Skip if no labels

# Create input function
pred_fn_new = tf.estimator.inputs.pandas_input_fn(
    x=x_new,
    batch_size=len(x_new),
    shuffle=False
)

Step 4: Generate Predictions and Evaluate

# Get predictions from the model
predictions = list(model.predict(input_fn=pred_fn_new))
final_preds = [pred['class_ids'][0] for pred in predictions]

# If you have labels for the new data, evaluate metrics
from sklearn.metrics import classification_report
print(classification_report(y_new, final_preds))

If you're doing unlabeled inference (no labels in the new data), just skip the evaluation step—you'll still get the predicted class IDs from the predict output.

内容的提问来源于stack exchange,提问作者Michael Yadidya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:25:50