Spark未指定路径却读取Maven项目resources中Hadoop配置的原理咨询
resources Directory Great question! Let me walk you through exactly what's happening here—this is a combination of Maven's resource handling, Hadoop's native configuration logic, and Spark's integration with Hadoop.
First, Maven's role: The
src/main/resourcesdirectory in your Maven project is a special, default location for resource files. When you runmvn compile,mvn package, or even run your project directly via Maven, Maven automatically copies all files from this directory into the compiled output (eithertarget/classesfor unbundled classes, or the root of your built JAR file). Critically, these files end up on the JVM's classpath—the list of locations the JVM checks for resources and classes by default.Hadoop's configuration loading behavior: Under the hood, Spark relies heavily on Hadoop's client libraries. Hadoop's
org.apache.hadoop.conf.Configurationclass (which Spark uses for all Hadoop-related operations like HDFS access) has built-in logic: it automatically scans the JVM's classpath for standard Hadoop config files likecore-site.xml,hdfs-site.xml,yarn-site.xml, etc. It doesn't require you to specify a path for these files—if they're on the classpath, they get loaded by default.Spark inherits this behavior: When you initialize a
SparkSessionorSparkContext, Spark creates an instance of Hadoop'sConfigurationunder the hood. It doesn't override Hadoop's default config loading logic, so it picks up exactly the files that Hadoop would—including the ones you've packaged into your JAR's classpath from Maven'sresourcesdirectory.
A quick example to visualize
Your project structure looks something like this:
my-spark-project/ ├── src/ │ ├── main/ │ │ ├── scala/ │ │ │ └── MyHdfsTest.scala │ │ └── resources/ │ │ ├── core-site.xml │ │ └── hdfs-site.xml └── pom.xml
When you run mvn package, Maven builds a JAR where core-site.xml and hdfs-site.xml live at the root of the JAR. When you submit this JAR to a cluster (or run it locally), the JVM running your Spark code sees these files on the classpath, and Hadoop's Configuration loads them automatically.
One extra note on priority
If you're running Spark on a cluster that already has Hadoop configs, your packaged configs from the resources directory will take precedence over the cluster's default configs. That's why your project works reliably across different clusters—you're bundling exactly the configs you need, and they override any cluster-specific defaults.
If you ever did need to specify a custom path for configs, you could set the spark.hadoop.conf.dir system property or manually load configs in your code, but in your case, the default classpath-based loading is doing all the work for you.
内容的提问来源于stack exchange,提问作者udit

