OpenShift 3.9 Jenkins Pipeline构建JBoss镜像偶发失败排查请求
Let's break down your intermittent build failure issue and walk through actionable steps to diagnose and fix it. Your error Error from server (BadRequest): build jboss-2 encountered an error: No logs are available. suggests the build pod either failed to start properly or was terminated before logs could be captured—here's how to dig deeper:
First, Get the Real Failure Context
The Jenkins error message is vague because the build didn't generate logs. The first step is to check OpenShift's internal events for the failed build:
- List all builds to find the failed one:
oc get builds -n ${DEV_PROJECT} - Describe the failed build to see detailed events:
oc describe build jboss-2 -n ${DEV_PROJECT}
Look at the Events section—this will often reveal the root cause (e.g., out-of-memory kills, node resource shortages, or registry issues) that Jenkins didn't capture.
Common Causes & Fixes
1. Insufficient Resources for Build Pods
Intermittent failures often stem from resource contention on OpenShift nodes. If your build pod doesn't have explicit resource limits, it might get terminated when the node runs out of CPU/memory.
- Check current BuildConfig resource settings:
oc describe bc jboss -n ${DEV_PROJECT} - Add resource requests/limits to your BuildConfig to guarantee resources for the build:
Update your BuildConfig definition (or add flags when creating it) with:
Or when creating the BuildConfig initially:strategy: sourceStrategy: from: kind: ImageStreamTag name: jboss-eap70-openshift:1.5 resources: requests: cpu: "1" memory: "1Gi" limits: cpu: "2" memory: "2Gi"openshift.newBuild("--name=jboss", "--image-stream=jboss-eap70-openshift:1.5", "--binary=true", "--build-env BUILD_CPU_REQUEST=1", "--build-env BUILD_MEMORY_REQUEST=1Gi")
2. Binary Upload Timeouts
The --from-dir flag uploads your local directory to the OpenShift API server. Network latency or large directory sizes can cause intermittent timeouts.
- Increase request timeout for the
start-buildcommand:
Modify your pipeline code to add a longer timeout:openshift.selector("bc", "jboss").startBuild("--from-dir=oc-build", "--wait=true", "--request-timeout=300s") - Optimize upload size:
Compress your build directory before uploading, then add a build hook to unpack it during the build:
Then update your BuildConfig to include a post-build hook that unpacks the archive (add this to thestage('Build Image with app') { sh "rm -rf oc-build && mkdir -p oc-build/deployments" sh "cp /var/lib/jenkins/jobs/cicd/jobs/cicd-tasks-pipeline/workspace/target/hello-1.0.war oc-build/deployments/ROOT.war" sh "tar czf oc-build.tar.gz oc-build" // Compress the directory openshift.withCluster() { openshift.withProject(env.DEV_PROJECT) { openshift.selector("bc", "jboss").startBuild("--from-file=oc-build.tar.gz", "--wait=true", "--request-timeout=300s") } } }strategysection):strategy: sourceStrategy: from: kind: ImageStreamTag name: jboss-eap70-openshift:1.5 hooks: postCommit: command: ["tar", "xzf", "oc-build.tar.gz", "-C", "."]
3. Verify File Integrity in Jenkins Workspace
Occasionally, file copy operations can fail silently (e.g., due to filesystem locks or permissions). Add a checksum validation step to ensure your WAR file is copied correctly:
stage('Build Image with app') { sh "rm -rf oc-build && mkdir -p oc-build/deployments" sh "cp /var/lib/jenkins/jobs/cicd/jobs/cicd-tasks-pipeline/workspace/target/hello-1.0.war oc-build/deployments/ROOT.war" // Validate file integrity sh "md5sum /var/lib/jenkins/jobs/cicd/jobs/cicd-tasks-pipeline/workspace/target/hello-1.0.war > original.md5" sh "md5sum oc-build/deployments/ROOT.war > copied.md5" sh "diff original.md5 copied.md5 || (echo 'WAR file copy corrupted!' && exit 1)" // Proceed with build openshift.withCluster() { openshift.withProject(env.DEV_PROJECT) { openshift.selector("bc", "jboss").startBuild("--from-dir=oc-build", "--wait=true", "--request-timeout=300s") } } }
4. Add Retry Logic for Intermittent Failures
Since the issue is sporadic, adding a retry mechanism to your pipeline can mitigate transient failures:
stage('Build Image with app') { sh "rm -rf oc-build && mkdir -p oc-build/deployments" sh "cp /var/lib/jenkins/jobs/cicd/jobs/cicd-tasks-pipeline/workspace/target/hello-1.0.war oc-build/deployments/ROOT.war" openshift.withCluster() { openshift.withProject(env.DEV_PROJECT) { // Retry up to 2 times (total 3 attempts) retry(2) { openshift.selector("bc", "jboss").startBuild("--from-dir=oc-build", "--wait=true", "--request-timeout=300s") } } } }
5. Check OpenShift Registry & Node Health
- Verify the OpenShift image registry is running and has enough storage:
oc get pods -n openshift-image-registry oc exec -n openshift-image-registry <registry-pod-name> -- df -h /registry - Check node resource utilization to rule out node-level resource exhaustion:
oc describe nodes
Start with checking the build events via oc describe build—that will give you the most precise clue about what's going wrong. From there, apply the relevant fix based on the root cause you find.
内容的提问来源于stack exchange,提问作者sudhir

