Apache Superset通过Spark连接MariaDB的配置问题求助
Let's tackle each of your issues one by one to get your MariaDB→Spark→Superset pipeline up and running smoothly:
1. Fixing pip3 install mysqlclient failures
This error almost always comes from missing system-level dependencies needed to compile the mysqlclient binary. Here's how to resolve it based on your operating system:
Ubuntu/Debian:
Install the required development packages first:sudo apt-get update && sudo apt-get install python3-dev default-libmysqlclient-dev build-essentialThen retry the pip install:
pip3 install mysqlclientCentOS/RHEL:
Install the corresponding dev packages:sudo yum install python3-devel mysql-devel gccThen run the pip install command again.
If you still run into issues, use the pure-Python mysql-connector-python as an alternative:
pip3 install mysql-connector-python
When pairing this with SQLAlchemy, your connection URI will look like: mysql+mysqlconnector://user:password@host:port/dbname
2. Accessing MariaDB data in Spark/Spark SQL
Your current hive-site.xml config is for setting up a Hive metastore stored in MariaDB, not for reading MariaDB data directly as a source. Let's fix this:
First, correct your spark-defaults.conf (remove spaces around the equals sign—Spark configs are strict about this formatting):
spark.driver.extraClassPath=/usr/share/java/mysql-connector-java.jar spark.executor.extraClassPath=/usr/share/java/mysql-connector-java.jar
Restart Spark services after updating this file.
To make MariaDB tables accessible in Spark SQL, create a mapped persistent table using the JDBC connector:
CREATE TABLE mariadb_your_table USING org.apache.spark.sql.jdbc OPTIONS ( url "jdbc:mysql://localhost:3306/your_db_name", dbtable "your_table_name", user "your_username", password "your_password", driver "com.mysql.jdbc.Driver" );
After running this, you can query the table directly with SELECT * FROM mariadb_your_table; in Spark SQL. Your Scala code works, but creating a persistent table makes it reusable across sessions.
3. Adding SQLAlchemy dialects in Superset
You don't need to modify Superset's core code to add dialects—you can configure this in your superset_config.py:
First install the required dialect packages (e.g., for PyHive/Spark):
pip3 install pyhive thriftAdd this to your
superset_config.pyto register the Hive/Spark dialect:from sqlalchemy.dialects import registry # Register Hive dialect for Spark Thrift Server registry.register("hive", "pyhive.sqlalchemy_hive", "HiveDialect")
Restart Superset after updating the config, and the dialect will be available when adding new data sources.
4. Correcting Superset's database URI configuration
You're mixing up two critical configurations:
SQLALCHEMY_DATABASE_URIinsuperset_config.pyis for Superset's own internal metadata database (the one that stores dashboards, users, etc.), not for connecting to Spark.
To connect to Spark in Superset:
- Start the Spark Thrift Server first (default port is 10000, not 4040—4040 is the Spark UI port):
$SPARK_HOME/sbin/start-thriftserver.sh - In the Superset UI, go to Data > Databases > + Database
- Choose "Hive" as the database type, then enter the connection string:
Or use the JDBC URI if preferred:hive://localhost:10000/jdbc:hive2://localhost:10000/
5. Fixing CSV import errors
Since you haven't shared the exact error message, here are common troubleshooting steps:
- File permissions: Ensure the CSV file is readable by the user running Superset/Spark. If it's on a remote filesystem, confirm the path is properly mounted and accessible.
- CSV format issues: When importing, explicitly set the delimiter, header presence, and encoding. For example, in Spark SQL:
CREATE TABLE csv_imported_table USING csv OPTIONS ( path "/absolute/path/to/your.csv", header "true", inferSchema "true", delimiter ",", encoding "UTF-8" ); - Superset temp directory: If importing via the Superset UI, check that the temp directory (configured via
UPLOAD_FOLDERinsuperset_config.py) has write permissions for the Superset process.
内容的提问来源于stack exchange,提问作者Code_Help

