OpenShift Enterprise 3.11部署deploy_cluster.yml失败及TLS握手超时求助
Let’s walk through step-by-step fixes for your problem—since your prerequisites.yml ran successfully, we can rule out basic environment setup issues, but there are a few key areas to dig into:
1. Verify Network & Port Access (The #1 Cause of TLS Timeouts)
TLS handshake failures almost always come down to blocked traffic or unreachable endpoints:
- Check critical OpenShift ports: Ensure ports 6443 (API server), 8443 (web console), and 2379/2380 (etcd) are open on your master nodes. Use
firewall-cmd --list-allto confirm, or temporarily disable firewalld for testing (not recommended for production):systemctl stop firewalld systemctl disable firewalld - SELinux enforcement: If SELinux is in
Enforcingmode, it might block container network traffic. Check withgetenforce, and temporarily set it toPermissiveto test:setenforce 0 - Test endpoint reachability: On your master node, run
telnet localhost 6443—if it fails, the API server isn’t listening on that port. From worker nodes, test connectivity to the master’s 6443 port too.
2. Fix Docker & Kubelet Service Issues
Your docker ps shows no running containers, which means the OpenShift control plane containers never started. Let’s check the underlying services:
- Check Docker status: Verify Docker is running without errors:
Look for issues like failed image pulls or daemon configuration errors. If you see problems, restart Docker:systemctl status docker journalctl -u docker -fsystemctl restart docker - Validate Kubelet health: Kubelet is responsible for starting OpenShift’s containers. Check if it’s running:
Common issues here include misconfiguredsystemctl status kubelet journalctl -u kubelet -fkubeconfigfiles or missing dependencies. If Kubelet is down, start it withsystemctl start kubelet.
3. Dig Into Ansible Deployment Logs
The "wait for control plane pods to appear" error means the deployment couldn’t detect the API server or etcd pods. Re-run the deployment with verbose logging to get granular details:
ansible-playbook -i playbooks/deploy_cluster.yml -vv
Look for failures before the wait step—common culprits include:
- Failed certificate generation for the API server
- Etcd cluster initialization failures
- Insufficient CPU/memory resources on master nodes
4. Check TLS Certificate Validity
Invalid or missing TLS certificates will break API server connectivity:
- Inspect master certificates: Go to
/etc/origin/master/and verify files likeapiserver.crt,apiserver.key, andca.crtexist with correct permissions (root:root, 600):ls -l /etc/origin/master/*.crt /etc/origin/master/*.key - Validate certificate details: Ensure the API server certificate includes your master’s hostname and IP in its Subject Alternative Names (SANs):
openssl x509 -in /etc/origin/master/apiserver.crt -text -noout | grep -A 10 "Subject Alternative Name" - Regenerate certificates if needed: If certificates are invalid, re-run the master cert configuration playbook:
ansible-playbook -i playbooks/openshift-master/config.yml --tags=master-certs
5. Correct oc login Syntax (Quick Fix!)
You mentioned using oc login -u sysadmin:admin—that’s likely a typo. The correct system admin user syntax is:
oc login -u system:admin
Or if you created an admin user earlier, use oc login -u admin (no colon needed).
Final Steps
After fixing the above issues:
- Re-run the deployment playbook:
ansible-playbook -i playbooks/deploy_cluster.yml - Once complete, check control plane pods with:
oc get pods -n openshift-kube-apiserver - Test
oc loginagain to confirm connectivity.
内容的提问来源于stack exchange,提问作者Prithiviraj Pandurangan

