如何将Jupyter Notebook Scala内核与Apache Spark集成?已装内核但遇问题
Hey there, let's get your Spark running smoothly in the Jupyter Scala kernel! First, let's recap your setup to make sure we're aligned:
Your Current Environment
- You've successfully installed the Scala kernel via jupyter-scala, confirmed with
jupyter kernelspec list:python3 /usr/local/homebrew/Cellar/python3/3.6.4_2/Frameworks/Python.framework/Versions/3.6/lib/python3.6/site-packages/ipykernel/resources scala /Users/bobyfarell/Library/Jupyter/kernels/scala - You're hitting errors when initializing Spark with code like:
val sparkHome = "/opt/spark-2.3.0-bin-hadoop2.7" val scalaVersion = scala.util.Properties.versionNumberString // (presumably followed by SparkContext/SparkSession setup)
The root issue here is that the Jupyter Scala kernel (built on Ammonite) doesn't include Spark dependencies out of the box. Here are three straightforward fixes to resolve this:
1. Load Spark Dependencies Directly in Your Notebook
This is the fastest way to test things out. Add these lines at the very top of your notebook to pull in Spark's core libraries (match the version to your local Spark install):
// Pull in Spark 2.3.0 dependencies via Ivy import $ivy.`org.apache.spark::spark-core:2.3.0` import $ivy.`org.apache.spark::spark-sql:2.3.0` // Initialize SparkSession import org.apache.spark.sql.SparkSession val spark = SparkSession.builder() .master("local[*]") // Use all local CPU cores for testing .appName("JupyterSparkDemo") .getOrCreate() // Verify it works with a simple test spark.sparkContext.parallelize(1 to 10).count()
If you need other Spark modules (like streaming or MLlib), just add more import $ivy. lines with the corresponding artifact (e.g., import $ivy.org.apache.spark::spark-streaming:2.3.0``).
2. Configure the Scala Kernel to Load Spark By Default
If you don't want to add the imports every time, modify the kernel's config file to include Spark's classpath permanently:
- Open the kernel config file at
/Users/bobyfarell/Library/Jupyter/kernels/scala/kernel.json - Update the
argvsection to include Spark's jar directory. Your config might look like this:{ "argv": [ "/usr/local/bin/ammonite", // Replace with your actual Ammonite path (check your original config) "--class-path", "/opt/spark-2.3.0-bin-hadoop2.7/jars/*", "--jupyter", "--connection-file", "{connection_file}" ], "display_name": "Scala", "language": "scala" } - Save the file and restart Jupyter Notebook. Now every Scala kernel session will automatically load Spark's dependencies.
3. Ensure Spark Environment Variables Are Set
Sometimes Spark needs explicit environment variables to locate its native libraries and config files. Check if SPARK_HOME is set in your notebook:
// Check if SPARK_HOME is defined println(System.getenv("SPARK_HOME")) // If it's empty, set it manually System.setProperty("spark.home", "/opt/spark-2.3.0-bin-hadoop2.7")
This ensures Spark can access all its required resources correctly.
If you run into specific error messages (like class-not-found exceptions), share those details and we can troubleshoot further!
内容的提问来源于stack exchange,提问作者Dimon Buzz

