You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Flink high-availability.storageDir维护规范及自动清理配置咨询

回答:

Let's tackle your questions one by one, based on hands-on experience with Flink's Zookeeper-based high availability setup:


1. 清理high-availability.storageDir的正确方式

Yes, you can safely delete old files/directories except for the default/blob folder—here's the breakdown:

  • The default/blob directory is Flink's distributed blob store, holding critical data like checkpoint metadata, job JARs, and user artifacts. Never delete this unless you're 100% sure all associated jobs are stopped and will never need to be restored.
  • All other subdirectories under /opt/flink/ha/ are typically leftover HA metadata from stopped jobs or temporary state directories from old JobManagers. These are no longer used once the job is terminated, so they're safe to remove.
  • Practical cleanup tips:
    • First, verify that the jobs linked to the target directories are fully stopped (no running JobManager/TaskManager processes referencing them).
    • Use a command like this to batch-clean old directories (adjust the -mtime value to match your retention needs):
      find /opt/flink/ha/ -type d -not -path "*/default/blob/*" -mtime +7 -delete
      
    • Avoid deleting directories while jobs are running—this can cause unexpected JobManager failures.

2. Automatic HA storage rotation configurations in flink-conf.yaml

Flink does have built-in settings to automate HA storage maintenance, so you don't have to clean manually. Add these parameters to your config:

  • Job history archiving & cleanup:
    • jobmanager.archive.fs.dir: Define a dedicated directory for archiving job history data. When a job stops, Flink will move its HA metadata here, and old HA directories can be auto-cleaned.
    • jobmanager.archive.fs.cleanup.interval: How often Flink runs cleanup for the archive directory (e.g., 1d for daily).
    • jobmanager.archive.fs.retention-time: How long archived data is kept (e.g., 7d to retain for 7 days).
  • Checkpoint auto-cleanup:
    • state.backend.fs.cleanup-interval: Interval for cleaning up expired checkpoints (e.g., 1h).
    • state.checkpoints.num-retained: Number of recent checkpoints to keep (e.g., 3—older ones will be deleted automatically).

These settings will prevent disk and inode exhaustion by automatically pruning unneeded data.


3. Fixing JobManager startup failures due to Zookeeper stale data

When disk exhaustion caused your JobManager to fail to start, deleting ZK data worked because:

  • Flink stores leader election data, job metadata, and HA state in Zookeeper. If the JobManager crashes abruptly (due to full disk), it might leave stale entries in ZK that block new JobManagers from taking over.
  • Preventive steps:
    • Enable the auto-cleanup settings above to avoid disk full scenarios in the first place.
    • Monitor disk usage and inode counts with alerts (e.g., trigger an alert when disk usage hits 80%).
    • If you run into this issue again, use Flink's built-in ZK client to clean stale job data safely:
      ./bin/zookeeper-client.sh deleteall /flink/ha/jobmanager/<your-job-id>
      
    Just make sure the job is fully stopped before running this command.

内容的提问来源于stack exchange,提问作者hdjur_jcv

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 09:04:26