You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Windows独立PySpark脚本中使用Mongo-Spark Connector并克隆仓库

How to Use Mongo-Spark Connector with Standalone PySpark Scripts on Windows

Hey there! Let's walk through how to get the Mongo-Spark Connector set up for your standalone PySpark script on Windows—no fancy jargon, just straightforward steps tailored for a Python beginner.

1. Clone the Mongo-Spark Connector Repository

First, make sure you have Git installed on your Windows machine (if not, grab Git for Windows—it's free and simple to set up). Open Git Bash or Command Prompt, then run this command to clone the repository:

git clone https://github.com/mongodb/mongo-spark.git

Once cloning finishes, navigate into the repository folder:

cd mongo-spark

2. Compile the Connector (If You Need the Latest/Custom Version)

If you want to use the absolute latest code or make custom tweaks, you'll need to compile the connector. Here's what you need first:

  • Java JDK 8 or 11 (Spark works best with these versions)
  • Scala (match the version your Spark uses—for example, Spark 3.3.x pairs with Scala 2.12)
  • Maven (to handle the build process)

Once those are installed, run this command in the repository folder to compile (skipping tests to speed things up):

mvn clean package -DskipTests

After compilation, you'll find the built jar file in the target subfolder—look for something like mongo-spark-connector_2.12-<version-number>.jar.

3. Use the Connector in Your PySpark Script

You have two easy ways to hook up the connector to your standalone PySpark script:

Option 1: Pass the Jar When Launching PySpark

Open Command Prompt or PowerShell, and start PySpark with the --jars flag pointing to your compiled connector jar (plus the required Mongo Java Driver jar—you can download this from Maven Central or find it in your local Maven repo at C:\Users\<YourUsername>\.m2\repository\org\mongodb\mongo-java-driver\<version>\):

pyspark --jars "C:\path\to\mongo-spark-connector_2.12-<version>.jar,C:\path\to\mongo-java-driver-<version>.jar"

Option 2: Configure SparkSession Directly in Your Script

If you're writing a standalone .py file (like mongo_spark_demo.py), you can set up the SparkSession to include the jars and Mongo connection details right in your code:

from pyspark.sql import SparkSession

# Initialize SparkSession with connector jars and Mongo configs
spark = SparkSession.builder \
    .appName("MongoToSparkDemo") \
    .config("spark.jars", "C:\\path\\to\\mongo-spark-connector_2.12-<version>.jar,C:\\path\\to\\mongo-java-driver-<version>.jar") \
    .config("spark.mongodb.input.uri", "mongodb://localhost:27017/your_database.your_collection") \
    .config("spark.mongodb.output.uri", "mongodb://localhost:27017/your_database.your_collection") \
    .getOrCreate()

# Load data from your Mongo collection into a Spark DataFrame
mongo_df = spark.read.format("mongodb").load()

# Check out your data!
mongo_df.show()

Then run your script with:

python mongo_spark_demo.py

4. Pro Tip for Beginners: Skip Cloning (Use Pre-Built Jars)

You don't actually need to clone the repository if you just want a working connector. Head to Maven Central, grab the pre-compiled jar that matches your Spark and Scala versions, and follow the same steps above to include it in your script way faster.

内容的提问来源于stack exchange,提问作者Vikrant Sonawane

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:53:33