Windows Service进入空闲状态问题及持续运行方案咨询
问题分析与解决方案
嘿,我帮你分析下这个问题——你的Windows服务运行几小时后就“躺平”不干活,重启又能正常跑,结合你贴的代码来看,核心问题大概率出在定时器的异常处理缺失或者定时器的使用方式上。下面是几个能让服务持续稳定运行的可行方案:
1. 给定时器回调加全局异常捕获,确保定时器能重启
你当前的代码里,如果PerformOtranOperations方法中任何一步抛出未捕获的异常(比如数据库连接超时、CRM接口调用失败),后面的mainTimer.Start()就不会执行,定时器直接停摆,服务看起来还在运行,但再也不会触发任务了。
修改代码,把所有逻辑包裹在try/catch/finally里,确保无论出什么错,定时器都能重新启动:
public void PerformOtranOperations(object sender, EventArgs e) { try { eventID++; eventLog1.WriteEntry(DateTime.Now.ToString("MM/dd/yyyy hh:mm:ss") + " - checking for new otran records", EventLogEntryType.Information, eventID); // 检查新记录 int otranRecords = _otranBL.GetOtranRecordCount(eventLog1); if (otranRecords == 0) { eventLog1.WriteEntry("0 new otran records", EventLogEntryType.Information, eventID); return; } eventLog1.WriteEntry(otranRecords.ToString("N0") + " new otran records found with proc_status = 0", EventLogEntryType.Information, eventID); // 处理记录 eventLog1.WriteEntry(DateTime.Now.ToString("MM/dd/yyyy hh:mm:ss") + " - Begin processing new otran records", EventLogEntryType.Information, eventID); int processedrecords = _otranBL.ProcessNewOtranRecords(eventLog1); eventLog1.WriteEntry(processedrecords.ToString() + " processed records", EventLogEntryType.Information, eventID); } catch (Exception ex) { eventID++; // 把错误日志写清楚,方便排查问题 eventLog1.WriteEntry($"Error during Otran processing: {ex.Message}\nStack Trace: {ex.StackTrace}", EventLogEntryType.Error, eventID); } finally { // 无论成功失败,都确保定时器重新启动 if (mainTimer != null && !mainTimer.Enabled) { mainTimer.Start(); } } }
2. 换成更稳定的System.Threading.Timer
System.Timers.Timer在跨线程异常场景下容易出现静默失效的问题,而System.Threading.Timer是专门为后台服务设计的,稳定性更好,而且不需要手动重启定时器——它会自动按间隔触发任务,哪怕某次回调出错,下次依然会执行。
修改你的定时器初始化和回调逻辑:
// 替换定时器类型 private System.Threading.Timer mainTimer; protected override void OnStart(string[] args) { eventLog1.WriteEntry("Service Start"); // 初始化:30秒后第一次执行,之后每30秒执行一次 mainTimer = new System.Threading.Timer(PerformOtranOperations, null, TimeSpan.FromSeconds(30), TimeSpan.FromSeconds(30)); } protected override void OnStop() { eventLog1.WriteEntry("Service Stopped"); // 先停止定时器,再释放资源 mainTimer?.Change(Timeout.Infinite, Timeout.Infinite); mainTimer?.Dispose(); mainTimer = null; } // 调整回调方法签名,适配System.Threading.Timer private void PerformOtranOperations(object state) { try { eventID++; eventLog1.WriteEntry(DateTime.Now.ToString("MM/dd/yyyy hh:mm:ss") + " - checking for new otran records", EventLogEntryType.Information, eventID); // 原有的业务逻辑不变 int otranRecords = _otranBL.GetOtranRecordCount(eventLog1); if (otranRecords == 0) { eventLog1.WriteEntry("0 new otran records", EventLogEntryType.Information, eventID); return; } eventLog1.WriteEntry(otranRecords.ToString("N0") + " new otran records found with proc_status = 0", EventLogEntryType.Information, eventID); eventLog1.WriteEntry(DateTime.Now.ToString("MM/dd/yyyy hh:mm:ss") + " - Begin processing new otran records", EventLogEntryType.Information, eventID); int processedrecords = _otranBL.ProcessNewOtranRecords(eventLog1); eventLog1.WriteEntry(processedrecords.ToString() + " processed records", EventLogEntryType.Information, eventID); } catch (Exception ex) { eventID++; eventLog1.WriteEntry($"Error during Otran processing: {ex.Message}\nStack Trace: {ex.StackTrace}", EventLogEntryType.Error, eventID); } }
3. 给服务加自动恢复机制
除了代码层面的优化,还可以利用Windows服务的内置故障恢复功能,让系统在服务异常时自动重启:
- 打开服务管理器,找到你的
CrmProcessOtran服务 - 右键点击→属性→切换到“恢复”选项卡
- 设置第一次、第二次、后续失败时的操作为“重启服务”,并设置重启前的等待时间(比如1分钟)
- 勾选“重置失败计数为0”的时间间隔(比如1天)
这样就算服务真的因为严重问题挂掉,系统也能自动帮你重启恢复。
4. 检查数据库连接是否泄漏
长时间运行后服务失效,还有可能是数据库连接池耗尽导致的。你要确保OtranBL里的数据库连接都用using语句包裹,这样连接会被自动释放回连接池:
// 举个例子,在OtranBL的方法里这么写 public int GetOtranRecordCount(EventLog eventLog) { // 使用using确保连接自动释放 using (var conn = new SqlConnection("你的数据库连接字符串")) { conn.Open(); // 执行查询逻辑,比如调用存储过程 using (var cmd = new SqlCommand("你的存储过程名", conn)) { cmd.CommandType = CommandType.StoredProcedure; return (int)cmd.ExecuteScalar(); } } }
内容的提问来源于stack exchange,提问作者Ram
相关产品推荐
相关产品推荐

