通过分离因变量与自变量实现基于神经网络的污染时序预测
Got it, let's walk through how to tackle this pollution concentration time series prediction task using neural networks. Here's a structured, practical approach tailored to your dataset:
First, let's organize your raw data into a clean table for clarity:
| 日期 | 污染浓度 | 露点 | 气温 | 气压 | 风向 | 风速 | 降雪量 | 降雨量 |
|---|---|---|---|---|---|---|---|---|
| 2010-01-02 00:00:00 | 129.0 | -16 | -4.0 | 1020.0 | SE | 1.79 | 0 | 0 |
| 2010-01-02 01:00:00 | 148.0 | -15 | -4.0 | 1020.0 | SE | 2.68 | 0 | 0 |
| 2010-01-02 02:00:00 | 159.0 | -11 | -5.0 | 1021.0 | SE | 3.57 | 0 | 0 |
| 2010-01-02 03:00:00 | 181.0 | -7 | -5.0 | 1022.0 | SE | 5.36 | 1 | 0 |
| 2010-01-02 04:00:00 | 138.0 | -7 | -5.0 | 1022.0 | SE | 6.25 | 2 | 0 |
Variable Split:
- Dependent Variable (Target):
污染浓度(this is what we’re predicting) - Independent Variables (Features):
露点,气温,气压,风向,风速,降雪量,降雨量(plus time-based features we’ll create later)
Neural networks perform best with properly formatted data—here’s what you need to do first:
- Encode Categorical Features: The
风向column is categorical (e.g., SE). Use one-hot encoding instead of label encoding to avoid introducing artificial order bias. For example, "SE" becomes a binary column where 1 indicates SE, 0 otherwise. - Engineer Time Features: Extract hour of day, day of week, etc., from the
日期column—time series models rely on these periodic patterns to make accurate predictions. - Normalize Numeric Features: Scale all numeric features (露点, 气温, 气压, 风速, 降雪量, 降雨量) to a range like [0,1] using
MinMaxScaler(from scikit-learn). This helps the neural network converge faster and more stably. - Create Sequences: For time series prediction, structure your data into input-output sequences. For example, use features from hours t-3, t-2, t-1 to predict the pollution concentration at hour t.
For time series tasks, LSTM (Long Short-Term Memory) networks are ideal because they excel at capturing temporal dependencies in sequential data. Here’s a basic implementation using TensorFlow/Keras:
import numpy as np from sklearn.preprocessing import MinMaxScaler from tensorflow.keras.models import Sequential from tensorflow.keras.layers import LSTM, Dense # Assume you've preprocessed data into sequences: # X_train/X_test: shape (num_samples, sequence_length, num_features) # y_train/y_test: shape (num_samples, 1) (target pollution concentration) # Build the LSTM model model = Sequential([ LSTM(50, return_sequences=True, input_shape=(X_train.shape[1], X_train.shape[2])), LSTM(50), Dense(1) # Output layer: predicts a single pollution concentration value ]) # Compile the model model.compile(optimizer='adam', loss='mean_squared_error') # Train the model history = model.fit( X_train, y_train, epochs=50, batch_size=32, validation_data=(X_test, y_test), verbose=1 )
After training, evaluate your model using regression-specific metrics:
- RMSE (Root Mean Squared Error): Measures average prediction error in the same units as your target (pollution concentration)
- MAE (Mean Absolute Error): Intuitive metric for average absolute prediction error
Also, plot training vs. validation loss to check for overfitting—if validation loss starts rising while training loss drops, add dropout layers or reduce the number of epochs to fix it.
内容的提问来源于stack exchange,提问作者Roman

