Azure DevOps Pipeline预配消息发送InternalServerError随机报错
问题背景
为排查相关异常,按如下路径新建流水线:New Pipeline > Azure Repos Git > Start pipeline,未对基础模板的流水线代码做任何修改,默认模板代码如下:
# Starter pipeline # Start with a minimal pipeline that you can customize to build and deploy your code. # Add steps that build, run tests, deploy, and more: # https://aka.ms/yaml trigger: - main pool: vmImage: ubuntu-latest steps: - script: echo Hello, world! displayName: 'Run a one-line script' - script: | echo Add other tasks to build, test, and deploy your project. echo See https://aka.ms/yaml displayName: 'Run a multi-line script'
流水线运行过程中会随机(每15次运行至少出现1次)弹出如下警告:
There was a failure in sending the provision message: Unexpected response code from remote provider InternalServerError
由于使用的是未做任何修改的官方入门模板,可排除YAML配置导致问题的可能,异常来源指向Azure DevOps服务或托管代理侧。
开启诊断日志后,在Agent_20220611-100355-utc.log中发现如下报错:
[2022-06-11 10:03:56Z ERR JobNotification] Connection to monitor port 49100 failed! [2022-06-11 10:03:56Z ERR JobNotification] System.Net.Internals.SocketExceptionFactory+ExtendedSocketException (111): Connection refused [::ffff:127.0.0.1]:49100 at System.Net.Sockets.Socket.DoConnect(EndPoint endPointSnapshot, SocketAddress socketAddress) at System.Net.Sockets.Socket.Connect(EndPoint remoteEP) at System.Net.Sockets.Socket.Connect(IPAddress address, Int32 port) at Microsoft.VisualStudio.Services.Agent.JobNotification.ConnectMonitor(String monitorSocketAddress) [2022-06-11 10:03:56Z ERR JobNotification] Invalid socket address hello. Job Notification will be disabled.
可行排查与解决建议
- 先确认影响范围:该报错属于代理内置的作业状态通知组件初始化失败,日志明确提示
Job Notification will be disabled,不会影响流水线核心步骤的执行,如果流水线最终运行状态为成功,仅弹出该警告,可暂时忽略,无需额外处理。 - 添加自动重试规则覆盖瞬时故障:托管代理运行在Azure共享资源池,这类随机出现的
InternalServerError属于平台侧资源调度的瞬时故障,可在流水线作业配置中添加失败重试策略,出现异常时自动重试1-2次即可规避绝大多数偶发问题,配置参考:
jobs: - job: defaultJob pool: vmImage: ubuntu-latest # 任务失败时自动重试2次 retryCountOnTaskFailure: 2 steps: # 原有流水线步骤
- 切换代理镜像版本验证:当前使用的
ubuntu-latest为滚动更新的镜像标签,可临时切换为固定版本的镜像(如ubuntu-22.04、ubuntu-20.04)运行一段时间,观察警告是否复现,排除特定镜像版本内置代理组件的兼容性问题。 - 调整代理规格降低资源抢占概率:默认共享托管代理的资源池租户较多,容易出现资源抢占、调度超时问题,可切换为更高规格的托管代理实例,降低供给失败的概率。
- 收集信息提交平台支持:如果该警告出现频率持续升高,甚至直接导致流水线运行失败,可收集出现异常的流水线运行ID、触发时间、完整诊断日志,提交给Azure DevOps官方支持团队排查资源供给链路的底层故障,这类平台侧内部错误无法通过终端用户配置完全修复,需要平台侧调整调度逻辑。
- 临时替代方案:如果该偶发问题已经严重影响流水线使用,可搭建自托管代理池运行流水线,自托管代理不需要经过平台侧托管实例的调度、供给流程,不会触发这类资源供给阶段的报错。
内容的提问来源于stack exchange,提问作者Jeffrey
相关产品推荐
相关产品推荐

