咨询:核SVM对比线性SVM在扩展特征空间中的劣势答题错误原因
Let’s walk through the critical issues with each of your points—even though some might sound intuitive at first glance, they contain fundamental conceptual misunderstandings that likely led to the full 0 score:
Problem with Point 1
Your statement: "若数据在扩展特征空间线性可分,线性SVM的间隔最大化效果更好且能得到更稀疏的解"
- The core mistake here is misframing kernel SVM and linear SVM as distinct methods in the expanded feature space. Kernel SVM is mathematically equivalent to a linear SVM trained on data explicitly mapped to the expanded space—it just uses the kernel trick to avoid the computational cost of explicit mapping. If the data is linearly separable in the expanded space, both approaches will find the exact same maximum-margin hyperplane. There’s no "better" performance or sparser solution for the linear SVM here; your characterization treats them as unrelated methods, which is incorrect.
Problem with Point 2
Your statement: "面对大数据集时,线性SVM的训练与预测耗时少于扩展特征空间中的核SVM"
- First, "扩展特征空间中的核SVM" is a misnomer. The entire purpose of the kernel trick is to avoid explicitly expanding the feature space—kernel SVM never operates on the expanded data directly. If you were to explicitly map data to a high-dimensional space and train a linear SVM, that’s a separate (and computationally costly) approach, not kernel SVM.
- While it’s true that kernel SVM has higher computational complexity (typically O(n³) for training, vs. more scalable O(n²) or better for optimized linear SVM solvers), your incorrect description of how kernel SVM works invalidates this point entirely. The professor likely saw this as a failure to grasp the kernel trick’s core purpose.
Problem with Point 3
Your statement: "核SVM相比线性SVM更易过拟合,生成更复杂的模型"
- This is an overly broad and unsubstantiated claim. Overfitting in SVMs depends almost entirely on the regularization parameter (
C) and kernel choice—not whether it’s a kernel or linear SVM. A linear SVM with a very highC(weak regularization) can overfit just as easily as a poorly tuned kernel SVM. Additionally, kernel SVMs don’t inherently produce "more complex models"—a kernel SVM using a linear kernel is identical to a standard linear SVM. Your generalization ignores the role of regularization and kernel capacity, making it an imprecise and incorrect statement for a technical exam.
Key Takeaway
Across all three points, the root issue is a lack of precise understanding of kernel SVM’s core mechanism: it is a linear SVM operating in a high-dimensional space, enabled by the kernel trick to avoid explicit data mapping. Your answers either misrepresented how kernel methods work, compared unrelated approaches, or made ungrounded generalizations—all critical errors in a machine learning exam setting.
内容的提问来源于stack exchange,提问作者Sushirdeep

