Spark读取MongoDB失败求助:基于HDP Sandbox 2.6环境
Hey there, let's work through this MongoDB-Spark connection issue you're hitting in Zeppelin on your HDP Sandbox 2.6 setup. I've debugged similar compatibility and configuration headaches before, so let's break down the fixes step by step:
You've added the correct version-matched jars (mongo-spark-connector_2.11-2.2.0 and mongo-java-driver 3.5.0 for Spark 2.2.0/Scala 2.11.8) to spark2-client/jars, but there are two extra critical checks here:
- Remove conflicting jars: Make sure there are no older versions of MongoDB connectors/drivers in the same directory—these can cause silent classpath conflicts that throw vague errors.
- Zeppelin interpreter access: Zeppelin's Spark interpreter might not automatically pick up jars from the Spark client directory. Copy those two jars to Zeppelin's Spark interpreter folder (typically
/usr/hdp/current/zeppelin-server/interpreter/spark) or add their paths to the interpreter's classpath settings (we'll cover that next).
Your current code cuts off, but the most common issue here is missing critical connection parameters in ReadConfig. You need to specify the full MongoDB URI (host, port, database, collection) along with your read preference. Here's a complete, working example:
import com.mongodb.spark.MongoSpark import com.mongodb.spark.config.ReadConfig // Configure full MongoDB connection details val readConfig = ReadConfig(Map( "spark.mongodb.input.uri" -> "mongodb://<your-mongo-host>:27017/<your-db>.<your-collection>", "readPreference.name" -> "primaryPreferred" // Adjust based on your cluster needs )) // If MongoDB requires authentication, use a URI like this instead: // "mongodb://<username>:<password>@<host>:27017/<db>.<collection>?authSource=admin" // Load the data into a Spark DataFrame val mongoDF = MongoSpark.load(spark.sqlContext, readConfig) // Test the connection with a simple preview mongoDF.show(5)
Zeppelin's Spark interpreter often needs explicit configuration to load external jars. Here's how to set it up:
- Open Zeppelin's UI, go to the Interpreters page, and find the
sparkinterpreter. - Scroll to the dependencies section, and add the full paths to your MongoDB jars (e.g.,
/usr/hdp/current/spark2-client/jars/mongo-spark-connector_2.11-2.2.0.jarand/usr/hdp/current/spark2-client/jars/mongo-java-driver-3.5.0.jar). - Alternatively, add these paths to the
spark.driver.extraClassPathandspark.executor.extraClassPathproperties in the interpreter's settings:spark.driver.extraClassPath=/usr/hdp/current/spark2-client/jars/mongo-spark-connector_2.11-2.2.0.jar:/usr/hdp/current/spark2-client/jars/mongo-java-driver-3.5.0.jar spark.executor.extraClassPath=/usr/hdp/current/spark2-client/jars/mongo-spark-connector_2.11-2.2.0.jar:/usr/hdp/current/spark2-client/jars/mongo-java-driver-3.5.0.jar - Save the settings and restart the Spark interpreter—this is non-negotiable for changes to take effect.
Sometimes the issue isn't code, but connectivity:
- From your HDP Sandbox, run
ping <your-mongo-host>to confirm network reachability. - Use
telnet <your-mongo-host> 27017(or your custom Mongo port) to check if the port is open. - Ensure your MongoDB instance isn't restricted to local access: Check the
bindIpsetting inmongod.conf—it should allow your Sandbox's IP, or be set to0.0.0.0for testing purposes. - If using MongoDB Atlas or a cloud instance, make sure your Sandbox's public IP is added to the MongoDB whitelist.
If you're still getting errors, enable debug logging to get a clearer picture of what's failing:
import org.apache.log4j.Logger import org.apache.log4j.Level // Enable debug logs for MongoDB Spark connector Logger.getLogger("org.mongodb.spark").setLevel(Level.DEBUG)
Rerun your code and check the Zeppelin Spark driver logs—this will show you specific errors like authentication failures, missing collections, or connection timeouts that the generic error message hides.
内容的提问来源于stack exchange,提问作者Chaouki

