在KOPS搭建的K8s集群部署Argo Workflow的疑问及流程咨询
Hey there! Let's break down all your questions about deploying Argo Workflows on your kops-provisioned K8s cluster (1 master, 2 workers, 1 load balancer) step by step.
1. Where Does Argo Workflows Deploy?
Short answer: Almost always on your Worker nodes—and that's intentional.
Kubernetes Master nodes typically have taints applied (like node-role.kubernetes.io/master:NoSchedule) to prevent user workloads from running on them, since masters handle critical cluster control plane tasks (API server, etcd, scheduler).
Argo's core components (the controller pod and UI server pod) don't include tolerations for these master taints by default. So the Kubernetes scheduler will automatically place them on your Worker nodes.
If you really needed to run Argo on the Master (not recommended for production), you'd have to add tolerations to the Argo deployment manifests—but this isn't best practice, as it could impact cluster stability.
2. Core Workflow of How Argo Works
Argo Workflows is built around Kubernetes Custom Resource Definitions (CRDs), which let you define complex, multi-step workflows as native Kubernetes objects. Here's the high-level breakdown:
- Argo Controller: A dedicated pod that runs in your cluster, constantly listening for new
WorkflowCRs. It manages the entire lifecycle of your workflows—scheduling tasks, handling retries, resolving dependencies, and updating statuses in real time. - Workflow CR: This is your workflow blueprint, written in YAML. It specifies tasks, their execution order (via linear
stepsor a directed acyclic graphDAG), container images to use, commands to run, data passing rules, and failure policies. - Task Pods: Each task in your workflow runs as a separate Kubernetes Pod on a Worker node. By default, Argo uses the
dockerexecutor (runs tasks directly in the pod), but you can also use alternatives likedocker-in-dockerfor building container images. - Data Sharing: Tasks can pass data between each other using shared volumes or Argo Artifacts (integrations with storage systems like S3, GCS, or local persistent volumes).
- Argo UI: A web interface that lets you visualize workflow progress, view pod logs, retry failed tasks, and browse historical workflow runs.
3. Step-by-Step Deployment to Your Cluster
Let's get Argo up and running. We'll use Helm for a flexible, maintainable installation (it handles dependencies like Postgres automatically):
Prerequisites
- Your kops cluster is fully operational, and you have
kubectlconfigured to access it. - Helm 3 is installed on your local machine.
Step 1: Set Up the Argo Namespace
First, we'll isolate Argo components in their own namespace to keep things organized:
kubectl create namespace argo
Step 2: Install Argo Workflows via Helm
- Add the official Argo Helm repository:
helm repo add argo https://argoproj.github.io/argo-helm helm repo update - Install Argo Workflows with Postgres (to persist workflow history—critical for tracking past runs):
helm install argo-workflows argo/argo-workflows --namespace argo --set postgresql.enabled=true
If you prefer using raw Kubernetes manifests instead, download the stable quick-start Postgres manifest from the Argo Workflows GitHub repository, then apply it with:
kubectl apply -n argo -f ./your-downloaded-manifest.yaml
Step 3: Verify the Deployment
Check that all Argo-related pods are running successfully:
kubectl get pods -n argo
You should see pods for:
argo-workflows-controller-xxxx: The core workflow controllerargo-workflows-server-xxxx: The UI serverargo-workflows-postgresql-xxxx: The database for storing workflow history
Step 4: Access the Argo UI
Option 1: Local Port Forwarding (For Testing)
Run this command to forward traffic from your local machine to the Argo UI server:
kubectl port-forward svc/argo-workflows-server -n argo 2746:2746
Then open http://localhost:2746 in your browser to access the UI.
Option 2: Use Your Cluster's Load Balancer
If you want to make the UI accessible across your network, update the argo-workflows-server service to use the LoadBalancer type:
kubectl patch svc argo-workflows-server -n argo -p '{"spec": {"type": "LoadBalancer"}}'
Wait a minute or two for the load balancer to assign an external IP, then run this command to get it:
kubectl get svc argo-workflows-server -n argo
Use the external IP with port 2746 (e.g., http://<external-ip>:2746) to access the UI.
Step 5 (Optional): Install the Argo CLI
The Argo CLI makes it easier to submit, monitor, and manage workflows from your terminal. Download the latest stable binary for your operating system from the Argo Workflows GitHub releases page, then set it up:
# Example for Linux: chmod +x argo-linux-amd64 sudo mv argo-linux-amd64 /usr/local/bin/argo # Verify the installation argo version
4. Full Functional Flow of an Argo Workflow
Once Argo is set up, here's what happens when you run a workflow:
- Define Your Workflow: Write a YAML file (e.g.,
my-first-workflow.yaml) that outlines your tasks. For example, a simple two-step workflow that prints a greeting then runs a data processing script. - Submit the Workflow: Use the CLI or
kubectlto submit the workflow to your cluster:argo submit -n argo ./my-first-workflow.yaml - Controller Takes Over: The Argo Controller detects the new
WorkflowCR, parses the task dependencies, and tells the Kubernetes scheduler to run the first task's pod on a Worker node. - Task Execution: The task pod spins up, executes its specified commands, and reports its status (success/failure) back to the controller.
- Handle Dependencies: If your workflow uses sequential steps, the controller waits for the first task to complete before scheduling the next. For DAG workflows, it runs parallel tasks where possible to optimize execution time.
- Data Passing: If tasks need to share data (e.g., a CSV file generated by task 1 used by task 2), Argo uses either shared volumes or artifacts to transfer the data securely.
- Workflow Completion: Once all tasks finish, the workflow's status is updated to
Succeeded(orFailedif any task failed). The full workflow history is stored in Postgres for future reference. - Monitor & Debug: Use the Argo UI or CLI to check task logs, retry failed tasks, or delete completed workflows to free up cluster resources.
内容的提问来源于stack exchange,提问作者user9182249

