如何在OCI中通过PySpark直接连接对象存储?含Jupyter场景需求
Hey there! I get you looking for a more straightforward way to access OCI Object Storage with PySpark in Jupyter Notebook—turns out you don’t have to go through the HDFS connector detour. Oracle provides a dedicated PySpark connector for OCI Object Storage that lets you directly interact with your data, which is way more convenient. Let me walk you through how to set this up and share some examples:
Prerequisites
- Make sure you have PySpark installed in your Jupyter environment, plus the OCI Python SDK (
ocipackage—install viapip install ociif missing) - Have your OCI authentication ready: either a configured
~/.oci/configfile with API keys, or use Instance Principal if your Jupyter runs on an OCI compute instance - Grab the latest OCI Object Storage Spark connector JAR (compatible with your Spark version) — you can reference it directly via Maven when initializing Spark, no need to download manually
Jupyter Notebook Setup & Code Examples
1. Initialize SparkSession with Connector Config
Instead of configuring HDFS first, set up your SparkSession to use the OCI Spark connector directly:
from pyspark.sql import SparkSession # Initialize Spark with OCI connector spark = SparkSession.builder \ .appName("OCI-ObjectStorage-PySpark") \ # Replace with the latest connector version compatible with your Spark version .config("spark.jars.packages", "com.oracle.oci.sdk:oci-hadoop-spark:3.25.0") \ # Use API key auth (switch to "instance_principal" if running on OCI compute) .config("spark.hadoop.fs.oci.client.auth.type", "api_key") \ .config("spark.hadoop.fs.oci.client.config.file", "/home/your-user/.oci/config") \ .config("spark.hadoop.fs.oci.client.profile.name", "DEFAULT") \ .getOrCreate()
2. Read Data from OCI Object Storage
Use the oci://<bucket-name>@<tenancy-namespace>/<file-path> path format to read your data. For example, reading a CSV:
# Read a CSV file from OCI bucket df = spark.read.csv( "oci://my-data-bucket@my-oci-tenancy/raw-data/sales_data.csv", header=True, inferSchema=True ) # Preview the data df.show(5)
3. Write Data to OCI Object Storage
You can write DataFrames back to Object Storage in formats like Parquet, CSV, or JSON:
# Write DataFrame as Parquet (overwrite existing data if present) df.write.parquet( "oci://my-data-bucket@my-oci-tenancy/processed-data/sales_summary.parquet", mode="overwrite" )
Key Notes
- Connector Version: Always use a connector version that matches your Spark major version (e.g., Spark 3.3.x works with recent connector releases)
- Authentication Options: If using Instance Principal (for OCI-hosted Jupyter), just set
spark.hadoop.fs.oci.client.auth.typetoinstance_principaland skip the config file parameters - Permissions: Ensure your OCI user/instance has the right IAM policies (e.g.,
Allow any-user to read objects in bucket my-data-bucket where request.principal.type = 'Instance'for Instance Principal)
Alternative: Simplified HDFS Connector Usage (If You Prefer)
If you still want to use the HDFS connector approach, you can configure Spark to recognize OCI paths as HDFS paths by setting the HDFS connector parameters in your SparkSession. But this adds extra steps compared to the dedicated Spark connector, so the direct method is strongly recommended.
内容的提问来源于stack exchange,提问作者nithin

