You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Kafka Broker日志目录存在却意外关闭的问题排查与解决

Kafka Broker Error: "Shutdown broker because all log dirs in /tmp/kafka-logs have failed"

Let's break down this error, its possible causes (even when your log directory exists), and how to fix or avoid it.

What This Error Means

This error from Kafka's LogManager tells you that the broker can't access or use all of its configured log directories. Since Kafka relies entirely on these directories to store topic data, manage offsets, and maintain cluster state, it has no choice but to shut down to prevent data corruption or loss.

Possible Causes (Even When the Directory Exists)

You've confirmed /tmp/kafka-logs is present and not empty, so we can rule out full directory deletion. Here are other likely issues:

  • Permission misconfiguration: The user running Kafka may have lost read/write access to the directory or its contents. For example, if you restarted Kafka under a different user, or system permissions were modified.
  • File system issues: The /tmp mount could be set to read-only, or the underlying disk may be full/inode-depleted (a common issue with temporary file systems).
  • Corrupted checkpoint files: The recovery-point-offset-checkpoint, log-start-offset-checkpoint, or replication-offset-checkpoint files in the log directory are damaged. Kafka needs these to recover state on startup.
  • Partial file deletion: System cleanup tools (like tmpwatch) might not have deleted the entire directory, but removed critical files within it that Kafka depends on.
  • File descriptor exhaustion: The Kafka process has hit its limit for open file descriptors, so it can't access the log files in the directory.

Solutions & Mitigations

1. Verify Directory Permissions

First, confirm the Kafka runtime user has full access to the log directory:

  1. Find the user running Kafka:
    ps aux | grep kafka
    
  2. Fix ownership and permissions (replace ec2-user with your Kafka user if different):
    sudo chown -R ec2-user:ec2-user /tmp/kafka-logs
    sudo chmod -R 755 /tmp/kafka-logs
    

2. Check File System Health

  • Check if /tmp is read-only:
    mount | grep /tmp
    
    If you see ro in the output, remount as read-write:
    sudo mount -o remount,rw /tmp
    
  • Check disk space and inodes:
    df -h /tmp  # Check free disk space
    df -i /tmp  # Check free inodes (critical for many small files)
    

3. Repair Corrupted Checkpoint Files

If checkpoint files are damaged, you can safely delete them (Kafka will rebuild them on startup, though you may lose some recent offset data—back them up first):

mkdir -p /tmp/kafka-checkpoint-backup
cp /tmp/kafka-logs/*.checkpoint /tmp/kafka-checkpoint-backup/
rm /tmp/kafka-logs/*.checkpoint

Restart Kafka and monitor if it starts successfully.

4. Move Logs to a Permanent Directory (Long-Term Fix)

Using /tmp for Kafka logs is risky because most Linux systems automatically clean temporary directories. To fix this permanently:

  1. Edit your Kafka server.properties file:
    # Replace /tmp/kafka-logs with a permanent path
    log.dirs=/var/lib/kafka/logs
    
  2. Create the new directory and set permissions:
    sudo mkdir -p /var/lib/kafka/logs
    sudo chown -R ec2-user:ec2-user /var/lib/kafka/logs
    sudo chmod -R 755 /var/lib/kafka/logs
    
  3. Copy existing logs to the new directory (optional, to retain data):
    cp -r /tmp/kafka-logs/* /var/lib/kafka/logs/
    
  4. Restart Kafka.

5. Increase File Descriptor Limits

Kafka needs many open file descriptors to handle topic logs. To increase the limit:

  1. Edit /etc/security/limits.conf and add:
    ec2-user soft nofile 65536
    ec2-user hard nofile 65536
    
  2. Log out and back in, then restart Kafka.

内容的提问来源于stack exchange,提问作者WestCoastProjects

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:14:25