如何在Windows独立PySpark脚本中使用Mongo-Spark Connector并克隆仓库
Hey there! Let's walk through how to get the Mongo-Spark Connector set up for your standalone PySpark script on Windows—no fancy jargon, just straightforward steps tailored for a Python beginner.
1. Clone the Mongo-Spark Connector Repository
First, make sure you have Git installed on your Windows machine (if not, grab Git for Windows—it's free and simple to set up). Open Git Bash or Command Prompt, then run this command to clone the repository:
git clone https://github.com/mongodb/mongo-spark.git
Once cloning finishes, navigate into the repository folder:
cd mongo-spark
2. Compile the Connector (If You Need the Latest/Custom Version)
If you want to use the absolute latest code or make custom tweaks, you'll need to compile the connector. Here's what you need first:
- Java JDK 8 or 11 (Spark works best with these versions)
- Scala (match the version your Spark uses—for example, Spark 3.3.x pairs with Scala 2.12)
- Maven (to handle the build process)
Once those are installed, run this command in the repository folder to compile (skipping tests to speed things up):
mvn clean package -DskipTests
After compilation, you'll find the built jar file in the target subfolder—look for something like mongo-spark-connector_2.12-<version-number>.jar.
3. Use the Connector in Your PySpark Script
You have two easy ways to hook up the connector to your standalone PySpark script:
Option 1: Pass the Jar When Launching PySpark
Open Command Prompt or PowerShell, and start PySpark with the --jars flag pointing to your compiled connector jar (plus the required Mongo Java Driver jar—you can download this from Maven Central or find it in your local Maven repo at C:\Users\<YourUsername>\.m2\repository\org\mongodb\mongo-java-driver\<version>\):
pyspark --jars "C:\path\to\mongo-spark-connector_2.12-<version>.jar,C:\path\to\mongo-java-driver-<version>.jar"
Option 2: Configure SparkSession Directly in Your Script
If you're writing a standalone .py file (like mongo_spark_demo.py), you can set up the SparkSession to include the jars and Mongo connection details right in your code:
from pyspark.sql import SparkSession # Initialize SparkSession with connector jars and Mongo configs spark = SparkSession.builder \ .appName("MongoToSparkDemo") \ .config("spark.jars", "C:\\path\\to\\mongo-spark-connector_2.12-<version>.jar,C:\\path\\to\\mongo-java-driver-<version>.jar") \ .config("spark.mongodb.input.uri", "mongodb://localhost:27017/your_database.your_collection") \ .config("spark.mongodb.output.uri", "mongodb://localhost:27017/your_database.your_collection") \ .getOrCreate() # Load data from your Mongo collection into a Spark DataFrame mongo_df = spark.read.format("mongodb").load() # Check out your data! mongo_df.show()
Then run your script with:
python mongo_spark_demo.py
4. Pro Tip for Beginners: Skip Cloning (Use Pre-Built Jars)
You don't actually need to clone the repository if you just want a working connector. Head to Maven Central, grab the pre-compiled jar that matches your Spark and Scala versions, and follow the same steps above to include it in your script way faster.
内容的提问来源于stack exchange,提问作者Vikrant Sonawane

