PySpark写入空PostgreSQL数据库报错No suitable driver求助
Let's walk through exactly what's causing this error and how to fix it—your issue is a mix of incorrect JDBC configuration and small code mistakes.
1. Fix the JDBC URL (The Root Cause of the Driver Error)
Your JDBC URL uses the wrong protocol for PostgreSQL, which is why Spark can't find the right driver. The correct prefix is jdbc:postgresql:// (note the extra ql), and the database name should be appended with a slash, not a semicolon.
Update your URL to this:
database = "postgres" jdbcUrl = f"jdbc:postgresql://localhost:5432/{database}"
2. Ensure the PostgreSQL JDBC Driver is Loaded
Even if you downloaded the driver, PySpark needs to know where it is and which class to use. Here are two reliable ways to handle this:
Option A: Pass the driver jar when running your script
When using spark-submit, include the --jars flag pointing to your PostgreSQL JDBC jar file (make sure you download a version compatible with your Spark/Java version):
spark-submit --jars /path/to/postgresql-42.6.0.jar your_etl_script.py
Option B: Configure the driver in your code
Add the driver option directly to your write configuration to explicitly specify the PostgreSQL driver class:
.option("driver", "org.postgresql.Driver")
3. Fix Other Small Code Mistakes
Wrong dbtable Value
In your first code snippet, you passed a DataFrame (df) to the dbtable option—this is invalid. The dbtable parameter expects a string that names your target table (and schema, if you want to use one).
Field Name Typo
You defined a schema field named second_columns (plural), but then tried to select second_column (singular) from json_df. Double-check that these column names match exactly to avoid missing column errors later.
Full Working Example for Empty Databases
Since your PostgreSQL instance is completely empty, we'll first create a schema (optional but good practice) and let Spark auto-create the table based on your DataFrame's schema. Make sure your PostgreSQL user has permissions to create schemas and tables!
from pyspark.sql import SparkSession # Initialize Spark Session (add the driver jar config if not using --jars flag) spark = SparkSession.builder \ .appName("WriteToEmptyPostgres") \ .config("spark.jars", "/path/to/postgresql-42.6.0.jar") # Remove if using --jars .getOrCreate() # Configuration user = "your_postgres_username" password = "your_postgres_password" database = "postgres" jdbcUrl = f"jdbc:postgresql://localhost:5432/{database}" target_schema = "news_schema" target_table = "json_data" # Step 1: Create the schema if it doesn't exist spark.read.format("jdbc") \ .option("url", jdbcUrl) \ .option("user", user) \ .option("password", password) \ .option("driver", "org.postgresql.Driver") \ .option("dbtable", f"(CREATE SCHEMA IF NOT EXISTS {target_schema}) AS temp") \ .load() # Step 2: Write your DataFrame to the target table json_df.select("first_column", "second_column") \ .write.format("jdbc") \ .mode("overwrite") \ .option("url", jdbcUrl) \ .option("user", user) \ .option("password", password) \ .option("driver", "org.postgresql.Driver") \ .option("dbtable", f"{target_schema}.{target_table}") \ .save()
Why This Works
- The corrected JDBC URL tells Spark to use the PostgreSQL-specific driver instead of looking for a generic "postgres" driver that doesn't exist.
- Explicitly setting the driver class removes any ambiguity about which JDBC driver to use.
- Creating the schema first ensures Spark has a valid namespace to create your table in.
- Using the correct
dbtablestring points Spark to the exact table you want to create/write to.
内容的提问来源于stack exchange,提问作者delalma

