为何Spark由Scala开发却无需安装Scala即可部署?
Great question! This is a super common point of confusion for anyone getting started with Spark, so let's break down the key reasons clearly:
Spark ships with its own bundled Scala runtime and dependencies
When you download a pre-built Spark distribution (or compile one yourself), the package includes all the necessary Scala libraries and runtime components right out of the box. Take a look inside thejarsfolder of your Spark installation—you'll see dozens of Scala-related.jarfiles (likescala-library-*.jar,scala-reflect-*.jar) that Spark uses to run. Tools likespark-submitorspark-shellautomatically load these bundled libraries, so there's no need for a system-wide Scala installation.Scala is a compile-time dependency, not a runtime requirement for Spark deployment
Spark is written in Scala, which means the Spark developers needed a Scala environment to build the framework from source. But once Spark is compiled, it produces Java bytecode (since Scala compiles to the same bytecode as Java) that runs directly on the JVM (Java Virtual Machine). All you need to run Spark is a working Java installation—Scala itself doesn't need to be present on the system at runtime.Development vs. deployment separation
If you're writing your own Spark jobs in Scala, you will need a Scala environment (like the Scala SDK, sbt, or Maven) during development to compile your code. But when it's time to deploy your job, you can package your code into a fat JAR (using tools likesbt assembly) that includes all your dependencies, or rely on Spark's cluster to provide the shared Scala libraries. Either way, the machines running your Spark cluster don't need Scala installed.
As a quick test: Download a pre-built Spark package, install Java, and run ./bin/spark-shell—it'll start up perfectly without you ever touching a Scala installer.
内容的提问来源于stack exchange,提问作者mill

