如何用kafka-connect-hdfs传输Kafka Avro数据至HDFS?运行失败求助
Let’s walk through the most common issues that might be blocking your HDFS sink connector from working, based on your configuration details:
1. Fix the HDFS URL Format
Your current hdfs.url=xxx.xxx.xxx.xxx:xxxx is missing the required protocol prefix. HDFS URLs need to start with hdfs:// for the connector to recognize them correctly. For example:
hdfs.url=hdfs://xxx.xxx.xxx.xxx:xxxx
Without this prefix, the connector won’t know how to establish a connection to your Hadoop cluster.
2. Replace ~ with Absolute Paths
The tilde (~) shorthand for your home directory often doesn’t get resolved correctly when running Kafka Connect as a service (or even as a background process). Swap out the relative paths in your config for full absolute paths:
# Instead of ~/logs logs.dir=/home/your-actual-username/logs # Instead of ~/topics topics.dir=/home/your-actual-username/topics
Make sure these directories exist and the user running Kafka Connect has read/write permissions for them.
3. Verify Hadoop Config Permissions
Your hadoop.conf.dir = /etc/hadoop/conf/ setting requires that the Kafka Connect process has read access to the files in this directory (like core-site.xml and hdfs-site.xml). If the user running Connect doesn’t have permission to read these files, the connector won’t be able to pull critical Hadoop cluster settings.
Run this command to check permissions (replace connect-user with the actual user running Kafka Connect):
sudo -u connect-user ls /etc/hadoop/conf/
If you get a permission denied error, adjust the directory/file permissions or run Connect as a user with appropriate access.
4. Confirm Environment Variable & Log Access
Even though you set LOG_DIR in .bash_profile, if you’re running Kafka Connect as a system service (e.g., via systemd), the shell environment variables might not be loaded. Your explicit logs.dir in the config should override this, but double-check that:
- The
logs.dirpath exists - The Connect user can write to it
- Check the logs in that directory for specific error messages (look for files like
connect.logorhdfs-sink.log). The logs will tell you exactly what’s failing—whether it’s a connection timeout, missing directory, or authentication issue.
5. Validate Connector Configuration
Before restarting the connector, double-check that all required fields are set correctly. For example, if your HDFS cluster uses Kerberos authentication, you’ll need additional configs (like kerberos.principal and kerberos.keytab) that aren’t in your current setup. If it’s a simple unsecure cluster, the fixes above should cover most cases.
Start by applying these changes, then check the logs again—they’ll be your best resource for narrowing down the exact issue.
内容的提问来源于stack exchange,提问作者Ambarish Nag

