基于时序数据的1D CNN能否从A、B推断比值特征C?是否需显式引入?
Great question—let’s break this down clearly, since you’re asking about both model capability and feature engineering best practices that come up often in time-series tasks.
Let’s split this into theoretical possibility and practical tradeoffs:
Theoretically, yes: A 1D CNN (or any deep learning model with sufficient capacity) can learn to approximate the A/B relationship. CNNs use convolutional filters to capture local feature interactions, and division is a mathematical operation that can be represented as a nonlinear transformation—especially when combined with activation functions like ReLU or sigmoid. Given enough training data and model capacity, the network can implicitly learn to compute this ratio from A and B.
Practically, explicit C might be better: That said, adding feature C directly can make your model’s job easier in most cases:
- If the A/B ratio is a known critical signal for your task, giving the model this feature reduces its learning burden. This is especially valuable with small datasets, where the model might not have enough examples to reliably learn the ratio on its own.
- It speeds up training convergence—you’re doing the arithmetic work upfront, so the model doesn’t have to waste capacity learning this specific relationship.
- The ratio might normalize A and B (e.g., if both scale with an underlying variable, A/B removes that scaling), which can stabilize training and help the model focus on more meaningful patterns.
While deep learning models are more robust to multicollinearity than traditional statistical models (like linear regression), there are still solid reasons to cut redundant, highly correlated features:
Wasted computational efficiency: Extra correlated features increase input dimensionality, which adds parameters to early convolutional layers and slows down training. Wasting model capacity on redundant signals leaves less room to learn complex, task-relevant patterns.
Unstable training: Even if you don’t care about interpretability, highly correlated features can cause unstable gradient updates. The model might oscillate between assigning weight to A vs. C (since they carry nearly identical information), slowing convergence or leading to suboptimal performance.
Poor interpretability: If you want to understand which features drive predictions, high correlation blurs the picture. You won’t be able to tell if the model relies on A, B, or their ratio—they’re all merged into one signal.
Increased overfitting risk: Redundant features can amplify noise. For example, if A has measurement error, adding C (derived from A) might spread that noise further, making the model more likely to overfit to training data idiosyncrasies.
If the A/B ratio is a known key signal for your task, explicitly adding it is usually a smart move (especially with small datasets). If you do add C, consider whether you need to keep both A/B and the original features—depending on your data size and task, dropping one pair can reduce redundancy without losing information. For large datasets with a powerful model, the network can learn the ratio on its own, but it’ll take more time and capacity to do so.
内容的提问来源于stack exchange,提问作者Richard Sun

