You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何Sklearn中RFECV/LogisticRegression运行后期核心使用率下降?

分析RFECV并行负载变化的原因

Hey there! Let's dig into why you're seeing that shift from full 6-core utilization to just 2 cores under heavy load with your RFECV jobs. Here are the most likely explanations, broken down clearly:

1. RFECV的阶段特性导致负载变化

Your RFECV workflow has two key phases, and their parallel behavior differs a lot:

  • Initial recursive elimination phase: When you start, there are plenty of features to evaluate. With n_jobs=5 and cv=10, scikit-learn uses joblib to spawn 5 parallel processes (each handling 2 cross-validation folds at a time). Combined with your other Notebook's n_jobs=1 process, that's 6 total single-threaded jobs—perfectly saturating your 6-core CPU, which matches your initial top output.
  • Final model fitting phase: Once RFECV narrows down to the optimal number of features (after iterating through step=3 feature removals), it runs a final fit on the entire training dataset (X_lab, y_lab) using that optimal feature subset. Since your LogisticRegression uses the default liblinear solver (which is single-threaded), this step only uses one core. Pair that with your other Notebook's ongoing single-threaded job, and you get exactly two cores pinned at ~100% utilization—exactly what you're seeing in the later top output.

2. Scikit-learn并行机制的细节

Scikit-learn relies on joblib for parallel execution, and it's designed to avoid unnecessary overhead. If your feature set shrinks enough that each cross-validation fold's training time becomes very short, joblib might not keep all 5 processes busy (process creation/destruction overhead starts to outweigh benefits). But this usually causes fluctuating CPU usage rather than a stable 2-core load, so this is less likely than the phase shift explanation.

3. 排除CPU过载/热节流的可能性

You mentioned a good water cooling setup and sufficient memory, which rules out thermal throttling as a cause. If throttling were happening, you'd see all cores reduce their utilization (and possibly CPU frequency drops) rather than just two cores running full tilt. You can confirm this with the sensors command to check CPU temps—they should stay well below your Intel i5-8400's max Tjunction of 100°C.


Quick Checks to Verify

  • Add print statements to track the current number of features during RFECV's fit:
    from sklearn.feature_selection import RFECV
    from sklearn.linear_model import LogisticRegression
    
    def print_feature_count(rfe, X, y):
        print(f"Current features remaining: {rfe.n_features_}")
        return rfe.fit(X, y)
    
    logreg = LogisticRegression()
    rfe = RFECV(logreg, step=3, cv=10, n_jobs=5)
    rfe = print_feature_count(rfe, X_lab, y_lab)
    
    When you see the feature count stop decreasing, that's when the final single-threaded fit starts.
  • Try switching your LogisticRegression solver to saga (which supports multi-threading for some operations) if you want to utilize more cores during the final fit:
    logreg = LogisticRegression(solver='saga', n_jobs=5)
    
    Note: saga works best with L1 regularization, so adjust your penalty parameter if needed.

内容的提问来源于stack exchange,提问作者Ted Callow

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 09:08:29