Keras深度学习模型输入维度缩减方法咨询
Hey Klaus, let's work through this input dimension bottleneck that's slowing down your travel time prediction model. From what you shared, that (12500,) input shape is almost certainly from one-hot encoding high-cardinality categorical features like your geographic coordinates—let's fix this step by step with practical, Keras-friendly solutions:
Your top priority should be replacing one-hot encoding for high-cardinality features (like unique geographic locations) with more efficient methods:
- Use Embedding Layers:
Map each categorical label (e.g., a location ID) to a low-dimensional dense vector (16-64 dimensions, adjust based on how many unique locations you have). This not only slashes dimension size but also captures meaningful relationships between locations (e.g., nearby areas might have similar embedding values). Here's a quick Keras example:from tensorflow.keras.layers import Input, Embedding, Flatten, Dense, concatenate from tensorflow.keras.models import Model # Input for geographic coordinates (integer-encoded location IDs) geo_input = Input(shape=(1,)) # Embed to 32-dimensional space geo_embedding = Embedding(input_dim=num_unique_locations, output_dim=32)(geo_input) geo_flat = Flatten()(geo_embedding) # Input for other 7 features (weather, time, distance, etc.) other_inputs = Input(shape=(7,)) combined_features = concatenate([geo_flat, other_inputs]) # Build regression model output = Dense(1)(combined_features) model = Model(inputs=[geo_input, other_inputs], outputs=output) - Target Encoding:
For categorical features, replace each category with the average travel time associated with it. This converts high-cardinality categories to a single numerical feature. Just make sure to use cross-validation to compute these averages (avoid leaking training data to validation sets) or add small noise to prevent overfitting.
Trim down your feature space by focusing on what matters:
- Filter Redundant Features:
Calculate correlation scores or mutual information between each feature and your target (travel time). Drop features that have near-zero correlation—they're just adding noise and increasing dimension size. For example, if you have both "hour of day" and "time of day (peak/off-peak)", you might only need one. - Normalize Numerical Features:
Features like distance should be scaled withStandardScalerorMinMaxScaler—this speeds up model convergence and prevents large-value features from dominating training. For time features, use cyclic encoding (e.g., sine/cosine transformations for hour of day) instead of one-hot encoding every hour.
Even with reduced dimensions, tweak your training pipeline to avoid slowdowns:
- Use a Lighter Model:
If you're using a deep fully-connected network, start with a smaller architecture (fewer neurons per layer) as a baseline. For example:from tensorflow.keras.models import Sequential model = Sequential([ Dense(64, activation='relu', input_shape=(reduced_feature_count,)), Dense(32, activation='relu'), Dense(1) # Regression output for travel time ]) - Speed Up Data Loading:
Switch totf.data.Datasetfor loading your data—it handles batching and prefetching to keep your GPU/CPU busy:import tensorflow as tf dataset = tf.data.Dataset.from_tensor_slices((X_processed, y)).batch(64).prefetch(tf.data.AUTOTUNE) model.fit(dataset, epochs=15) - Enable Mixed Precision Training:
Reduce memory usage and speed up training by using mixed float16 precision:from tensorflow.keras import mixed_precision mixed_precision.set_global_policy('mixed_float16')
If you still have a large feature space after optimizing encoding, try these:
- PCA (Principal Component Analysis):
Apply PCA to your numerical features (including encoded categorical features) to retain 95% of the variance. This automatically compresses the feature space while preserving most predictive information. Just remember to standardize your data first. - Avoid t-SNE for Training:
t-SNE is great for visualization, but it's computationally expensive and doesn't preserve predictive relationships well—stick to PCA for model input.
Start with fixing categorical encoding first—it's the most impactful fix for your dimension explosion issue, and it'll help your model learn better patterns than raw one-hot encoding.
内容的提问来源于stack exchange,提问作者Klaus

