You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Flink 1.18.1 TaskManager异常重启与Missing Resources日志解读求助

问题场景

使用Flink 1.18.1时,一条任务管线出现异常:TaskManager莫名重启且初始化受阻,仅抛出Java级别的致命错误:

# A fatal error has been detected by the Java Runtime Environment:
#  SIGSEGV (0xb) at pc=0x00007fc77a1944bf, pid=1, tid=391

JobManager中记录的相关日志如下:

2024-11-12 15:27:13.001 
2024-11-12 14:27:13,000 INFO  org.apache.flink.runtime.resourcemanager.slotmanager.DefaultSlotStatusSyncer [] - Starting allocation of slot 68f4b9a6f8243289ac95d2f32b7ec1cf from xx.xx.0.219:xxx-dd39bc for job f2e512b96b238a413633456f42b7ef21with resource profile ResourceProfile{cpuCores=1, taskHeapMemory=2.049gb (2199912448 bytes), taskOffHeapMemory=0 bytes, managedMemory=11.725gb (12589622671 bytes), networkMemory=2.726gb (2927204977 bytes)}.
2024-11-12 15:27:13.005 
2024-11-12 14:27:13,005 INFO  org.apache.flink.fs.s3.common.writer.S3Committer             [] - Committing checkpoint/app-name/rocksDBstorage/949294afcef3dfa56f0a6e2c7e4436f7/chk-1812/_metadata with MPU ID NWQyYTMxOGEtYzM1OS00MzMzLTkxYTEtZDdlNjQ3ZjgyNzI3LjI3N2Y0M2UzLTExMjMtNDI3Ny1hYWUzLWM2OTQ2NmVkNDA4NHgxNzMxNDIxNjMyOTU1MDk4MjQ5
2024-11-12 15:27:13.018 
2024-11-12 14:27:13,018 INFO  org.apache.flink.runtime.resourcemanager.slotmanager.DefaultSlotStatusSyncer [] - Freeing slot 68f4b9a6f8243289ac95d2f32b7ec1cf.
2024-11-12 15:27:13.019 
2024-11-12 14:27:13,019 INFO  org.apache.flink.runtime.jobmaster.JobMaster                 [] - Disconnect TaskExecutor xx.xx.0.219:xxx-dd39bc because: TaskExecutor pekko.tcp://flink@xx.xx.0.219:xxx/user/rpc/taskmanager_0 has no more allocated slots for job f2e512b96b238a413633456f42b7ef21.
2024-11-12 15:27:13.080 
2024-11-12 14:27:13,079 INFO  org.apache.flink.runtime.resourcemanager.slotmanager.FineGrainedSlotManager [] - Matching resource requirements against available resources.
2024-11-12 15:27:13.080 
Missing resources:

2024-11-12 15:27:13.080 
     Job f2e512b96b238a413633456f42b7ef21
2024-11-12 15:27:13.080 
        ResourceRequirement{resourceProfile=ResourceProfile{UNKNOWN}, numberOfRequiredSlots=1}
2024-11-12 15:27:13.080 
Current resources:
2024-11-12 15:27:13.080 
    TaskManager xx.xx.0.219:xxx-b6ea2e
2024-11-12 15:27:13.080 
        Available: ResourceProfile{cpuCores=0, taskHeapMemory=0 bytes, taskOffHeapMemory=0 bytes, managedMemory=0 bytes, networkMemory=0 bytes}
2024-11-12 15:27:13.080 
        Total:     ResourceProfile{cpuCores=1, taskHeapMemory=2.049gb (2199912448 bytes), taskOffHeapMemory=0 bytes, managedMemory=11.725gb (12589622671 bytes), networkMemory=2.726gb (2927204977 bytes)}
2024-11-12 15:27:13.080 
    TaskManager xx.xx.0.219:xxx-bbb4fa
2024-11-12 15:27:13.080 
        Available: ResourceProfile{cpuCores=0, taskHeapMemory=0 bytes, taskOffHeapMemory=0 bytes, managedMemory=0 bytes, networkMemory=0 bytes}
2024-11-12 15:27:13.080 
        Total:     ResourceProfile{cpuCores=1, taskHeapMemory=2.049gb (2199912448 bytes), taskOffHeapMemory=0 bytes, managedMemory=11.725gb (12589622671 bytes), networkMemory=2.726gb (2927204977 bytes)}
2024-11-12 15:27:13.080 
    TaskManager xx.xx.0.219:xxx-a1345b
2024-11-12 15:27:13.080 
        Available: ResourceProfile{cpuCores=0, taskHeapMemory=0 bytes, taskOffHeapMemory=0 bytes, managedMemory=0 bytes, networkMemory=0 bytes}
2024-11-12 15:27:13.080 
        Total:     ResourceProfile{cpuCores=1, taskHeapMemory=2.049gb (2199912448 bytes), taskOffHeapMemory=0 bytes, managedMemory=11.725gb (12589622671 bytes), networkMemory=2.726gb (2927204977 bytes)}

未找到"Missing Resources"日志的解读信息,该日志为INFO级别,请问这是管线资源不足导致的故障,还是Flink JobManager的正常行为,需要从其他方向排查故障原因?


解答

日志解读

"Missing Resources"这条INFO级别的日志是Flink SlotManager的正常调试输出,并非故障根源。日志中ResourceProfile{UNKNOWN}是核心线索:当TaskManager意外崩溃后,JobMaster无法获取该节点上任务的有效资源需求描述,因此生成UNKNOWN的资源配置,SlotManager会输出这条日志,它是TaskManager崩溃后的连锁反应,而非资源不足导致的问题。

故障根源定位

TaskManager抛出的SIGSEGV是JVM层面的致命内存错误,通常由以下原因引发:

  • 任务使用的RocksDB(从日志中可见状态存储基于RocksDB)Native代码崩溃
  • JVM版本与Flink 1.18.1不兼容
  • 操作系统层面内存不足,触发OOM killer
  • 第三方依赖的Native库冲突

排查建议

  • 找到TaskManager崩溃时生成的hs_err_pid*.log文件(通常在TaskManager工作目录下),该文件包含SIGSEGV的详细堆栈信息,可直接定位崩溃的具体模块
  • 检查RocksDB配置:确认managed memory分配未超出系统可用内存,关闭不必要的Native优化选项
  • 验证JVM版本是否符合Flink官方要求(推荐OpenJDK 11或17)
  • 查看TaskManager所在节点的系统日志,排查是否有硬件故障或OOM killer的执行记录

内容的提问来源于stack exchange,提问作者Baal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 10:10:58