You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在OCI中通过PySpark直接连接对象存储?含Jupyter场景需求

Hey there! I get you looking for a more straightforward way to access OCI Object Storage with PySpark in Jupyter Notebook—turns out you don’t have to go through the HDFS connector detour. Oracle provides a dedicated PySpark connector for OCI Object Storage that lets you directly interact with your data, which is way more convenient. Let me walk you through how to set this up and share some examples:

Direct PySpark Access to OCI Object Storage

Prerequisites

  • Make sure you have PySpark installed in your Jupyter environment, plus the OCI Python SDK (oci package—install via pip install oci if missing)
  • Have your OCI authentication ready: either a configured ~/.oci/config file with API keys, or use Instance Principal if your Jupyter runs on an OCI compute instance
  • Grab the latest OCI Object Storage Spark connector JAR (compatible with your Spark version) — you can reference it directly via Maven when initializing Spark, no need to download manually

Jupyter Notebook Setup & Code Examples

1. Initialize SparkSession with Connector Config

Instead of configuring HDFS first, set up your SparkSession to use the OCI Spark connector directly:

from pyspark.sql import SparkSession

# Initialize Spark with OCI connector
spark = SparkSession.builder \
    .appName("OCI-ObjectStorage-PySpark") \
    # Replace with the latest connector version compatible with your Spark version
    .config("spark.jars.packages", "com.oracle.oci.sdk:oci-hadoop-spark:3.25.0") \
    # Use API key auth (switch to "instance_principal" if running on OCI compute)
    .config("spark.hadoop.fs.oci.client.auth.type", "api_key") \
    .config("spark.hadoop.fs.oci.client.config.file", "/home/your-user/.oci/config") \
    .config("spark.hadoop.fs.oci.client.profile.name", "DEFAULT") \
    .getOrCreate()

2. Read Data from OCI Object Storage

Use the oci://<bucket-name>@<tenancy-namespace>/<file-path> path format to read your data. For example, reading a CSV:

# Read a CSV file from OCI bucket
df = spark.read.csv(
    "oci://my-data-bucket@my-oci-tenancy/raw-data/sales_data.csv",
    header=True,
    inferSchema=True
)

# Preview the data
df.show(5)

3. Write Data to OCI Object Storage

You can write DataFrames back to Object Storage in formats like Parquet, CSV, or JSON:

# Write DataFrame as Parquet (overwrite existing data if present)
df.write.parquet(
    "oci://my-data-bucket@my-oci-tenancy/processed-data/sales_summary.parquet",
    mode="overwrite"
)

Key Notes

  • Connector Version: Always use a connector version that matches your Spark major version (e.g., Spark 3.3.x works with recent connector releases)
  • Authentication Options: If using Instance Principal (for OCI-hosted Jupyter), just set spark.hadoop.fs.oci.client.auth.type to instance_principal and skip the config file parameters
  • Permissions: Ensure your OCI user/instance has the right IAM policies (e.g., Allow any-user to read objects in bucket my-data-bucket where request.principal.type = 'Instance' for Instance Principal)

Alternative: Simplified HDFS Connector Usage (If You Prefer)

If you still want to use the HDFS connector approach, you can configure Spark to recognize OCI paths as HDFS paths by setting the HDFS connector parameters in your SparkSession. But this adds extra steps compared to the dedicated Spark connector, so the direct method is strongly recommended.

内容的提问来源于stack exchange,提问作者nithin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 06:54:48