You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

构建Scala Spark JAR包时sbt assembly失败问题求助

Fixing sbt Assembly Failure for Your Spark Project

Hey there! Let's work through that sbt assembly failure you're hitting. Even though you didn't share the exact error message, based on your provided build.sbt and super common pitfalls with assembling Spark JARs, these are the most likely fixes to get you up and running:

1. Add the sbt-assembly Plugin (Critical!)

First off, sbt doesn't include the assembly plugin by default. You need to add it to your project's plugin configuration:

Create (or edit) the file project/plugins.sbt and add this line:

addSbtPlugin("com.eed3si9n" % "sbt-assembly" % "0.14.10")

(Note: This version is compatible with Scala 2.11.x and sbt 0.13.x, which matches your setup. If you're using a newer sbt version, you can pick a matching plugin version, but this one should work for your stack.)

2. Mark Spark Dependencies as provided

Spark clusters already include Spark core libraries by default. If you bundle these into your JAR, you'll end up with dependency conflicts (this is a top cause of assembly failures). Update your build.sbt to mark the Spark core dependency as provided:

scalaVersion := "2.11.8"
resolvers += "MavenRepository" at "http://central.maven.org/maven2"
resolvers += "spark-packages" at "https://dl.bintray.com/spark-packages/maven/"
libraryDependencies ++= {
  val sparkVersion = "2.2.1"
  Seq(
    "org.apache.spark" %% "spark-core" % sparkVersion % "provided"
  )
}

3. Configure Assembly Merge Strategy (For Duplicate Class Errors)

If you're seeing errors about duplicate files or classes (like META-INF files or conflicting dependency classes), you need to define a merge strategy in your build.sbt:

// Add this at the bottom of your build.sbt
assemblyMergeStrategy in assembly := {
  case PathList("META-INF", "MANIFEST.MF") => MergeStrategy.discard
  case PathList("META-INF", xs @ _*) => MergeStrategy.discard
  case x => MergeStrategy.first
}

This tells sbt to discard duplicate META-INF files (which are almost never needed in a Spark JAR) and pick the first version of any other conflicting files.

Full Updated build.sbt Example

Here's what your complete build.sbt should look like after applying the above fixes:

scalaVersion := "2.11.8"
resolvers += "MavenRepository" at "http://central.maven.org/maven2"
resolvers += "spark-packages" at "https://dl.bintray.com/spark-packages/maven/"
libraryDependencies ++= {
  val sparkVersion = "2.2.1"
  Seq(
    "org.apache.spark" %% "spark-core" % sparkVersion % "provided"
  )
}

assemblyMergeStrategy in assembly := {
  case PathList("META-INF", "MANIFEST.MF") => MergeStrategy.discard
  case PathList("META-INF", xs @ _*) => MergeStrategy.discard
  case x => MergeStrategy.first
}

If you still run into issues after trying these steps, share the exact error message you're getting (like duplicate class names or missing dependencies) and we can dive deeper!

内容的提问来源于stack exchange,提问作者syv

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:19:37