TFS 2018发布流程:"Deploy TestAgent"步骤引发服务器莫名重启
Hey Dave, sorry to hear you're hitting this intermittent headache after upgrading to TFS 2018! Let's walk through some targeted troubleshooting steps to fix that "OS shutting down/agent losing communication" error on Agent-19 when running the "Deploy TestAgent on localhost" step.
Likely Causes & Fixes
System Resource Crunches on Agent-19
Intermittent failures often pop up when a machine is stretched thin. Here's what to check:- Fire up Task Manager on Agent-19 and monitor memory, CPU, and disk usage during the deployment. If you see sustained high utilization (90%+), free up disk space, close unnecessary background apps/services, or consider upgrading the agent's hardware.
- Double-check that the TFS agent service runs with local admin privileges—TestAgent deployments need elevated access, and restricted permissions can lead to unexpected resource starvation or failures that mimic a shutdown.
Legacy TestAgent Tooling Compatibility
Since this worked flawlessly on TFS 2012, the deployment step might be using outdated installer or scripts that don't play nice with TFS 2018's runtime:- Update your TestAgent installer to match your TFS 2018 version (16.122.27102.1). Version mismatches between the server, agent, and TestAgent are a common source of intermittent issues.
- Enable verbose logging for the deployment step by adding
/l*v C:\TestAgentDeploy.logto the installer command in your pipeline. After a failure, dig into that log—you might spot hidden errors that trigger a system shutdown (unlikely, but worth ruling out).
Automated Shutdown Triggers on Agent-19
The error explicitly calls out the OS shutting down, so let's check for automated restarts:- Run
schtasks /queryin Command Prompt on Agent-19 to list all scheduled tasks. Look for any tasks that trigger a restart/shutdown around your deployment window—these could be conflicting with your pipeline. - Check Windows Update settings: if Agent-19 is set to auto-restart after updates, that's a prime suspect. Disable auto-restart temporarily or schedule updates outside your deployment times.
- Verify third-party management tools (like SCCM or monitoring software) aren't pushing restart commands to Agent-19 intermittently.
- Run
Agent-TFS Server Communication Glitches
Sometimes the "losing communication" error is a red herring for underlying connection issues:- Restart the
Azure DevOps Agentservice on Agent-19 (yes, that's the service name even for TFS 2018) to reset the communication channel with your TFS server. - Dig into the agent's diagnostic logs (default path:
C:\agent\_diag) around the time of failure. Look for entries about dropped connections or authentication errors—these can make the server think the agent's machine is shutting down.
- Restart the
Step Timeout Misconfiguration
TFS 2018 has stricter default timeouts compared to 2012. If the TestAgent deployment takes longer on Agent-19 than expected, the server might mark it as unresponsive:- Head into your release pipeline settings and increase the timeout for the "Deploy TestAgent on localhost" step. Give it a few extra minutes to complete—this could resolve false positives where the server thinks the agent is offline.
Quick Isolation Test
Try running the exact same TestAgent deployment command manually on Agent-19 multiple times. If it fails intermittently outside the TFS pipeline, the problem is definitely with the agent machine itself, not your TFS server setup.
内容的提问来源于stack exchange,提问作者Dave R

