如何让Docker容器获得稳定一致的执行时间
Docker容器内进程执行时间波动过大,无法满足稳定性要求
我需要用Docker隔离特定进程,该进程在多核虚拟机上重复运行多次,每次执行以挂钟时间测量记录,目标是将执行时间差控制在200ms以内,但Docker环境下最优与最差执行时间差约1秒,达不到要求。
附图表说明:蓝色柱为原生执行时间(稳定性良好),橙色柱为Docker进程执行时间(波动大)。
最小复现示例
1. mem.cpp(内存密集型C++程序)
#include <bits/stdc++.h> #include <vector> using namespace std; string CustomString(int len) { string result = ""; for (int i = 0; i<len; i++) result = result + 'm'; return result; } int main() { int len = 320; std::vector< string > arr; for (int i = 0; i < 100000; i++) { string s = CustomString(len); arr.push_back(s); } cout<<arr[10] <<"\n"; return 0; }
2. script.sh(Docker容器启动脚本,编译并运行程序,记录挂钟时间)
#!/bin/bash # compile the file g++ -O2 -std=c++17 -Wall -o _sol mem.cpp # execute file and record execution time (wall clock) ts=$(date +%s%N) ./_sol echo $((($(date +%s%N) - $ts)/1000000)) ms
3. Python并行执行脚本(通过ProcessPoolExecutor启动多个Docker容器执行任务)
import docker import logging import os import tarfile import tempfile from concurrent.futures import ProcessPoolExecutor log_format = '%(asctime)s %(threadName)s %(levelname)s: %(message)s' dkr = docker.from_env() def task(): ctr = dkr.containers.create("gcc:12-bullseye", command="/home/script.sh", working_dir="/home") # copy files into container cp_to_container(ctr, "./mem.cpp", "/home/mem.cpp") cp_to_container(ctr, "./script.sh", "/home/script.sh") # run container and capture logs ctr.start() ec = ctr.wait() logs = ctr.logs().decode() ctr.stop() ctr.remove() # handle error if (code := ec['StatusCode']) != 0: logging.error(f"Error occurred during execution with exit code {code}") logging.info(logs) def file_to_tar(src: str, fname: str): f = tempfile.NamedTemporaryFile() abs_src = os.path.abspath(src) with tarfile.open(fileobj=f, mode='w') as tar: tar.add(abs_src, arcname=fname, recursive=False) f.seek(0) return f def cp_to_container(ctr, src: str, dst: str): (dir, fname) = os.path.split(os.path.abspath(dst)) with file_to_tar(src, fname) as tar: ctr.put_archive(dir, tar) if __name__ == "__main__": # set logging level logging.basicConfig(level=logging.INFO, format=log_format) # start ProcessPoolExecutor ppex = ProcessPoolExecutor(max_workers=max(os.cpu_count()-1,1)) for _ in range(21): ppex.submit(task)
已尝试的优化措施
- 限制CPU核心数(8核虚拟机中使用4核及以下),执行时间波动无明显改善,推测问题可能出在Docker Engine层面。
补充发现
更换为新版gcc:13-bookworm镜像后,容器内执行时间比原生更优且稳定性大幅提升,推测原问题与镜像配置相关,希望彻底解决Docker容器内执行时间一致性问题。
内容的提问来源于stack exchange,提问作者Stefan Zhelyazkov
相关产品推荐
相关产品推荐

