Docker中基于H2O实现实时评分:仅初始化一次H2O的方案
问题:Docker中H2O模型实时评分的持久化运行方案
我是Docker新手,尝试通过Docker利用H2O模型实现实时评分。以下是我的Dockerfile、docker-compose.yml及代码:
Dockerfile
FROM python:3.9 WORKDIR /app RUN apt-get update && apt-get install -y default-jre RUN pip install h2o==3.38.0.3 RUN pip install pandas==1.3.5 COPY . /app #EXPOSE 8000 CMD ["python", "my_models.py"]
docker-compose.yml
version: '3' services: click_app: image: h2oai/h2o-open-source-k8s container_name: click_models restart: unless-stopped build: . ports: - "8001:8001"
核心代码
h2o.init(port=23023, nthreads=10) def main(): df_h2o = h2o.import_file('input.csv') df_h2o["PhoneNumber"] = df_h2o["PhoneNumber"].ascharacter() dfScore = df_h2o[df_h2o['NotDateMonth']==5] ms1gbm, ms1rf, ms1glm, ms2gbm, ms2rf, ms2glm = CallModels() predictions = Score(ms1gbm, ms1rf, ms1glm, ms2gbm, ms2rf, ms2glm, dfScore) print("NOTIFY...", predictions) return 0 if __name__ == '__main__': print(sys.argv) main()
当前问题:docker-compose.yml中设置restart: unless-stopped时,H2O会反复初始化,代码无限重复执行;移除该设置则仅运行一次。我希望仅初始化一次H2O并保持容器运行,有新数据集出现时再执行评分,请问有哪些实现方法?
解决方案
方案1:拆分H2O服务与评分逻辑
把H2O实例做成独立的长期运行服务,评分代码作为客户端连接远程H2O,实现一次初始化、多次调用:
- 修改docker-compose.yml,新增独立H2O服务:
version: '3' services: h2o_server: image: h2oai/h2o-open-source-k8s container_name: h2o_server restart: unless-stopped ports: - "23023:23023" command: ["java", "-jar", "/opt/h2o/h2o.jar", "-port", "23023", "-nthreads", "10"] scoring_app: build: . container_name: scoring_app restart: unless-stopped depends_on: - h2o_server volumes: - ./input_data:/app/input_data # 挂载数据目录用于监控新文件
- 修改评分代码,连接远程H2O并添加文件监控逻辑:
import h2o import time import os # 连接远程H2O服务,仅初始化一次 h2o.init(ip="h2o_server", port=23023) def process_file(file_path): df_h2o = h2o.import_file(file_path) df_h2o["PhoneNumber"] = df_h2o["PhoneNumber"].ascharacter() dfScore = df_h2o[df_h2o['NotDateMonth']==5] ms1gbm, ms1rf, ms1glm, ms2gbm, ms2rf, ms2glm = CallModels() predictions = Score(ms1gbm, ms1rf, ms1glm, ms2gbm, ms2rf, ms2glm, dfScore) print("NOTIFY...", predictions) def watch_and_score(): processed_files = set() data_dir = "/app/input_data" while True: current_files = set(os.listdir(data_dir)) new_files = current_files - processed_files for file in new_files: if file.endswith(".csv"): process_file(f"{data_dir}/{file}") processed_files.add(file) time.sleep(30) # 每30秒检查一次新文件 if __name__ == '__main__': watch_and_score()
方案2:将评分逻辑封装为API服务
用FastAPI搭建HTTP接口,容器长期运行API服务,收到请求时触发评分:
- 修改Dockerfile,添加API依赖:
FROM python:3.9 WORKDIR /app RUN apt-get update && apt-get install -y default-jre RUN pip install h2o==3.38.0.3 RUN pip install pandas==1.3.5 RUN pip install fastapi uvicorn COPY . /app EXPOSE 8001 CMD ["uvicorn", "my_models:app", "--host", "0.0.0.0", "--port", "8001"]
- 修改代码为API形式,提前初始化H2O和模型:
from fastapi import FastAPI, File, UploadFile import h2o import tempfile import os # 初始化一次H2O和模型 h2o.init(port=23023, nthreads=10) ms1gbm, ms1rf, ms1glm, ms2gbm, ms2rf, ms2glm = CallModels() app = FastAPI() @app.post("/score/local-file") def score_local_file(file_path: str): df_h2o = h2o.import_file(file_path) df_h2o["PhoneNumber"] = df_h2o["PhoneNumber"].ascharacter() dfScore = df_h2o[df_h2o['NotDateMonth']==5] predictions = Score(ms1gbm, ms1rf, ms1glm, ms2gbm, ms2rf, ms2glm, dfScore) return {"predictions": predictions.as_data_frame().to_dict()} @app.post("/score/upload") def score_uploaded_file(file: UploadFile = File(...)): with tempfile.NamedTemporaryFile(delete=False, suffix=".csv") as tmp: tmp.write(file.file.read()) tmp_path = tmp.name df_h2o = h2o.import_file(tmp_path) df_h2o["PhoneNumber"] = df_h2o["PhoneNumber"].ascharacter() dfScore = df_h2o[df_h2o['NotDateMonth']==5] predictions = Score(ms1gbm, ms1rf, ms1glm, ms2gbm, ms2rf, ms2glm, dfScore) os.unlink(tmp_path) return {"predictions": predictions.as_data_frame().to_dict()}
- 调整docker-compose.yml(去掉
image字段,挂载数据目录可选):
version: '3' services: click_app: container_name: click_models restart: unless-stopped build: . ports: - "8001:8001" volumes: - ./input_data:/app/input_data # 可选,用于访问本地文件路径
方案3:用文件监控工具触发评分
用inotifywait监控数据目录,容器长期运行,有新文件时自动执行评分脚本:
- 修改Dockerfile,安装监控工具并添加启动脚本:
FROM python:3.9 WORKDIR /app RUN apt-get update && apt-get install -y default-jre inotify-tools RUN pip install h2o==3.38.0.3 RUN pip install pandas==1.3.5 COPY . /app COPY start.sh /app/start.sh RUN chmod +x /app/start.sh CMD ["/app/start.sh"]
- 编写
start.sh启动脚本:
#!/bin/bash # 后台初始化H2O并保持运行 python -c "import h2o; h2o.init(port=23023, nthreads=10)" & H2O_PID=$! # 监控input目录下的csv文件创建/移动事件 inotifywait -m /app/input -e create -e moved_to | while read path action file; do if [[ $file == *.csv ]]; then echo "Processing new file: $file" python my_models.py "$path$file" fi done # 等待H2O进程,防止容器退出 wait $H2O_PID
- 修改评分代码,支持传入文件路径参数:
import sys import h2o # 初始化一次H2O h2o.init(port=23023, nthreads=10) def main(file_path): df_h2o = h2o.import_file(file_path) df_h2o["PhoneNumber"] = df_h2o["PhoneNumber"].ascharacter() dfScore = df_h2o[df_h2o['NotDateMonth']==5] ms1gbm, ms1rf, ms1glm, ms2gbm, ms2rf, ms2glm = CallModels() predictions = Score(ms1gbm, ms1rf, ms1glm, ms2gbm, ms2rf, ms2glm, dfScore) print("NOTIFY...", predictions) return 0 if __name__ == '__main__': if len(sys.argv) < 2: print("Please provide input file path") sys.exit(1) main(sys.argv[1])
- docker-compose.yml挂载数据目录:
version: '3' services: click_app: container_name: click_models restart: unless-stopped build: . ports: - "8001:8001" volumes: - ./input:/app/input
内容的提问来源于stack exchange,提问作者mlee_jordan
相关产品推荐
相关产品推荐

