如何在Windows 10中创建HDFS文件夹并搭建HDFS系统
Hey there! Let’s break down exactly how to create HDFS folders on Windows 10, set up a local HDFS instance, and connect it to a Cloudera Docker container. I’ve walked through these steps countless times, so let’s dive in.
Prerequisite: Set up winutils.exe
Windows doesn’t natively handle the system-level operations Hadoop needs, so winutils.exe is a must-have to avoid weird errors in Spark or HDFS commands. Here’s what to do:
- Grab the winutils package that matches your Hadoop/Spark version (e.g., if you’re using Spark 3.3, go with Hadoop 3.3.x).
- Extract it to a folder like
C:\hadoop\bin—make surewinutils.exeis directly in thatbindirectory. - Configure system environment variables:
- Create a new
HADOOP_HOMEvariable pointing toC:\hadoop(the parent folder ofbin). - Add
%HADOOP_HOME%\binto yourPathvariable.
- Create a new
- Test it: Open Command Prompt and run
winutils.exe ls /. If it doesn’t throw a "file not found" error, you’re good to go (it might complain about HDFS not running yet—don’t worry about that for now).
Set up a single-node HDFS instance on Windows 10
If you want to run HDFS locally on your Windows machine, follow these steps:
- Download Hadoop
Grab the Windows-compatible Hadoop binary package from Apache’s archives, then extract it toC:\hadoop. - Configure Hadoop files
Head toC:\hadoop\etc\hadoopand tweak these config files:hadoop-env.cmd: Addset JAVA_HOME=C:\Program Files\Java\jdk1.8.0_xxx(replace with your actual JDK path—if there are spaces, wrap it in quotes like"C:\Program Files\Java\jdk1.8.0_xxx").core-site.xml: Replace the existing content with this:<configuration> <property> <name>fs.defaultFS</name> <value>hdfs://localhost:9000</value> </property> <property> <name>hadoop.tmp.dir</name> <value>C:\hadoop\tmp</value> </property> </configuration>hdfs-site.xml: Use this config to set up single-node replication and skip permission checks (great for testing):<configuration> <property> <name>dfs.replication</name> <value>1</value> </property> <property> <name>dfs.permissions.enabled</name> <value>false</value> </property> </configuration>
- Format HDFS
Open Command Prompt, navigate toC:\hadoop\bin, and runhdfs namenode -format. You’ll see a "successfully formatted" message when it’s done—don’t run this more than once unless you want to wipe your HDFS data. - Start HDFS services
Runstart-dfs.cmdfrom thebinfolder. Two command windows will pop up (for namenode and datanode)—leave these open, they’re the HDFS services running in the background. - Verify it’s working
Open your browser and go tohttp://localhost:50070. If you see the HDFS management dashboard, you’re all set!
Two ways to create HDFS folders
Option 1: Use the HDFS command line
This is the quickest way for one-off tasks. Open Command Prompt and run:
hdfs dfs -mkdir /your-folder-name # Example: Create a folder named "raw-data" hdfs dfs -mkdir /raw-data
To confirm it was created, run:
hdfs dfs -ls /
Option 2: Create folders from a Spark app in IntelliJ
If you’re building a Spark application, you can create HDFS folders directly in your code. Here’s a Scala example:
import org.apache.hadoop.fs.{FileSystem, Path} import org.apache.spark.SparkContext import org.apache.spark.SparkConf object HdfsFolderCreator { def main(args: Array[String]): Unit = { val conf = new SparkConf() .setAppName("HdfsFolderCreator") .setMaster("local[*]") // Run locally for testing val sc = new SparkContext(conf) // Get a reference to the HDFS file system val hdfs = FileSystem.get(sc.hadoopConfiguration) val targetFolder = new Path("/spark-generated-folder") if (!hdfs.exists(targetFolder)) { hdfs.mkdirs(targetFolder) println("Folder created successfully!") } else { println("Folder already exists.") } sc.stop() } }
Make sure your Spark project includes the Hadoop client dependency. For Maven, add this to your pom.xml:
<dependency> <groupId>org.apache.hadoop</groupId> <artifactId>hadoop-client</artifactId> <version>3.3.4</version> <!-- Match your Hadoop version --> </dependency>
Connect to HDFS in a Cloudera Docker container
If you have a Cloudera Docker container running HDFS, here’s how to connect to it from Windows:
- Find the HDFS address of the container
SSH into the Cloudera container (or run a command inside it) and executehdfs getconf -confKey fs.defaultFS. You’ll get something likehdfs://cdh-master:8020. - Connect via command line
Run this in Command Prompt, replacing the address with your container’s HDFS URL:
Alternatively, you can update your localhdfs dfs -fs hdfs://cdh-master:8020 -mkdir /remote-project-foldercore-site.xml’sfs.defaultFSto point to this remote URL, then restart your local HDFS services (if you have them running). - Connect via Spark app
Update your SparkConf to point to the remote HDFS:val conf = new SparkConf() .setAppName("ConnectToRemoteHdfs") .setMaster("local[*]") .set("fs.defaultFS", "hdfs://cdh-master:8020") - Fix network issues
If Windows can’t resolve the container’s hostname (likecdh-master), edit yourC:\Windows\System32\drivers\etc\hostsfile and add a line like[container-IP] cdh-masterto map the hostname to the container’s IP address.
内容的提问来源于stack exchange,提问作者Janis Karklins

