如何用Python的hdfs3库上传本地文件至HDFS?求host/port获取方式
Hey there! I see you're trying to use hdfs3 to upload local files to HDFS with your Hue credentials—let's walk through how to get the HDFS host and port info you need for a successful connection.
There are a few reliable ways to get this info, and since you already have access to Hue, we'll start with the most convenient methods:
1. Grab Details Directly from Hue's UI
Since you're already using Hue, this is the quickest path:
- Log into your Hue dashboard.
- Navigate to Clusters (usually in the top menu bar) and select your HDFS cluster.
- Look for sections labeled NameNode Information or HDFS Configuration—here you'll find the NameNode host (this is your
hostvalue) and the RPC port (typically8020for managed clusters like Cloudera/Hortonworks, or9000for default Apache Hadoop setups). - If you don't see it right away, check Hue Settings → HDFS to view the configured NameNode address (it'll look like
host:port).
2. Fetch Configs via Hue Query
Since you know how to run queries in Hue, you can pull HDFS configuration details using Hive or Spark SQL:
Run this query to get the exact RPC address:
SELECT name, value FROM hdfs_configs WHERE name = 'dfs.namenode.rpc-address';
The result will give you a full host:port pair (e.g., namenode.yourcluster.com:8020). Split this into separate host and port values for your code.
3. Verify with Cluster Command Line (If You Have Access)
If you can SSH into any cluster node, run this command to get the RPC address directly:
hdfs getconf -confKey dfs.namenode.rpc-address
This will output the host and port in one line, which you can split for your script.
Once you have your host and port, here's how to structure your hdfs3 code (including authentication with your Hue credentials—adjust based on your cluster's auth setup):
from hdfs3 import HDFileSystem # Replace with your actual details hdfs = HDFileSystem( host='your-namenode-hostname', port=8020, # Update to your cluster's RPC port user='your-hue-username', password='your-hue-password' # Omit if your cluster uses Kerberos ) # Test the connection by listing your target upload path print(hdfs.ls('/your-confirmed-hdfs-upload-path')) # Upload your local file hdfs.put('/path/to/your/local/file.csv', '/your-confirmed-hdfs-upload-path/file.csv')
- The RPC port (not the HTTP port) is what hdfs3 uses for file operations—don't mix them up!
- If your cluster uses Kerberos instead of password auth, initialize a Kerberos ticket first and add
kerberos=Trueto theHDFileSystemconstructor instead of providing a password.
内容的提问来源于stack exchange,提问作者Ricky Zhou

