You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MLRun中Sklearn-classifier任务冻结(K8s Pod处于Pending状态)求助

MLRun中Sklearn-classifier任务冻结(Pod处于Pending状态)的排查与解决

问题场景

运行MLRun的Sklearn-classifier任务时出现冻结,任务持续数十分钟未结束,日志仅输出启动信息:

2023-02-21 13:50:15,853 [info] starting run training uid=e8e66defd91043dda62ae8b6795c74ea DB=http://mlrun-api:8080
2023-02-21 13:50:16,136 [info] Job is running in the background, pod: training-tgplm

Web UI显示对应Pod处于Pending状态(截图:Pod Pending状态)。

执行代码如下,调用classifier_fn.run(train_task, local=False)触发问题:

# Import the Sklearn classifier function from the function hub
classifier_fn = mlrun.import_function('hub://sklearn-classifier')

# Prepare the parameters list for the training function
training_params = {"model_name": ['risk_xgboost'],              
              "model_pkg_class": ['sklearn.ensemble.GradientBoostingClassifier']}

# Define the training task, including the feature vector, label and hyperparams definitions
train_task = mlrun.new_task('training', 
                      inputs={'dataset': transactions_fv.uri},
                      params={'label_column': 'n4_pd30'}
                     )

train_task.with_hyper_params(training_params, strategy='list', selector='max.accuracy')

# Specify the cluster image
classifier_fn.spec.image = 'mlrun/mlrun'

# Run training
classifier_fn.run(train_task, local=False)

排查与解决步骤

1. 检查集群资源是否充足

Pod处于Pending状态的核心原因通常是集群无足够CPU/内存分配:

  • 执行kubectl describe pod training-tgplm,查看Events字段的具体提示(如Insufficient cpu/Insufficient memory)
  • 若资源不足,可扩容集群节点,或调整任务资源请求:
    # 为任务指定资源请求与限制
    classifier_fn.with_requests(cpu="1", mem="2G")
    classifier_fn.with_limits(cpu="2", mem="4G")
    

2. 确认镜像拉取状态

指定的mlrun/mlrun通用镜像可能存在拉取失败或依赖缺失:

  • 通过kubectl describe pod training-tgplm检查Events,是否有ImagePullBackOff/ErrImagePull错误
  • 建议使用Sklearn-classifier官方适配的镜像(而非通用镜像):
    # 指定带版本的专用镜像,或直接使用hub函数默认镜像(无需手动设置)
    classifier_fn.spec.image = 'mlrun/mlrun:1.3.0'
    

3. 验证存储卷挂载可用性

任务输入的transactions_fv.uri对应的存储卷可能无法正常挂载:

  • 查看Pod描述中的Volumes和VolumeMounts部分,确认PVC、存储类存在且正常
  • 检查MLRun的artifact存储配置,确保集群能访问对应的存储服务(如S3、NFS)

4. 简化Hyper参数配置

当前单元素列表的Hyper参数配置虽语法合法,但可简化避免潜在解析问题:

# 改用单值而非列表
training_params = {"model_name": 'risk_xgboost',              
                   "model_pkg_class": 'sklearn.ensemble.GradientBoostingClassifier'}
train_task.with_hyper_params(training_params, strategy='list', selector='max.accuracy')

5. 检查MLRun服务连通性

确认mlrun-api服务正常运行且Pod可访问:

  • 执行kubectl get svc mlrun-api,确认服务状态为Running
  • 检查集群网络策略是否允许Pod访问mlrun-api的8080端口

内容的提问来源于stack exchange,提问作者JIST

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 20:57:04