Windows Server 2016中Service Control Manager频繁重启服务的原因排查
Got it, let's break down how to figure out why your service is getting restarted every minute by the Service Control Manager (SCM), triggered by Event ID 7031. Here's a step-by-step approach to diagnose the root cause:
Dig deeper into event logs beyond ID 7031
The 7031 event only tells you the service crashed and will restart—it doesn't explain why the crash happened. Head to Event Viewer → Windows Logs → System, then filter for related events that fire right before the 7031:- Event ID 7023: Will show a specific error message (e.g., "The service terminated with the following error: The system cannot find the file specified.") that directly points to the failure reason
- Event ID 7024: Includes an error code (like
0x80070005for permission issues) that narrows down the problem - Also check the Application Log for Event ID 1000 (Application Error)—this will reveal if the service's process crashed due to an unhandled exception, faulty DLL, or memory issue.
Verify the service's recovery configuration
Openservices.msc, find your target service, right-click → Properties → Recovery tab. Here you'll confirm that "Restart the service" is set for First/Second/Subsequent failures with a 60-second delay (which matches your issue). But this is just SCM's response—your real goal is fixing the underlying crash that triggers this loop. Also check the "Reset fail count after" value; if it's set too low, SCM might keep treating failures as consecutive.Check service dependencies
A service can crash if one of its required dependencies isn't running or fails. Go to the service's Properties → Dependencies tab. List all dependent services, then check each one's status and event logs. If a dependency is also crashing or failing to start, fixing that could resolve your main service's issues. Try manually starting all dependent services first, then launch your target service to see if it stays running.Analyze process dumps and service-specific logs
If the service process is crashing, generate a dump file to pinpoint the exact issue:- Open Task Manager → Details tab, locate the process linked to your service
- Right-click → Create dump file (it saves to a default path like
C:\Users\<YourUser>\AppData\Local\Temp) - Use tools like WinDbg (part of the Windows SDK) or Visual Studio to open the dump—look for exception codes, faulty modules, or stack traces that show where the crash occurred.
Also, check if the service has its own log files (usually in its installation directory, e.g.,C:\Program Files\<ServiceName>\Logs)—these often contain detailed error messages that system event logs miss.
Check system resources and service permissions
- Resources: Monitor the service's process in Task Manager for sudden CPU/memory spikes, which could indicate leaks or infinite loops causing termination. Low system-wide memory or disk space can also force services to crash.
- Permissions: Go to the service's Properties → Log On tab. If it runs under a custom account, verify that account has:
- Read/write access to the service's installation directory and configuration files
- Permissions to access required network resources (like shared folders or databases)
- Necessary user rights (check via Local Security Policy → Local Policies → User Rights Assignment)
As a temporary test, try setting the service to run under the Local System account to rule out permission issues (just note the security implications of this).
Update the service and system patches
Outdated software often has bugs that cause crashes. Check the service vendor's website for updates or hotfixes. Also, install the latest cumulative updates for Windows Server 2016—Microsoft frequently fixes SCM-related issues and compatibility bugs in these updates.
For example, if you see Event ID 7024 with error code 0x80070005, that's a permission denial—focus on fixing the service account's access to needed resources. If Event ID 1000 points to a missing/corrupted DLL, try re-registering the DLL or reinstalling the service.
内容的提问来源于stack exchange,提问作者Krithi B

