You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

SnappyData与Hive Metastore共存疑问:同一应用能否同时访问二者?

Can I Use Hive Metastore and SnappyData Metastore in the Same Spark Application?

Good question! You absolutely can use both Hive Metastore and SnappyData Metastore in the same application—you just need to adjust how you initialize your sessions or reference tables explicitly. Let’s break down the issue and fix your code:

Why Your Current Code Is Behaving This Way

When you add hive-site.xml and call .enableHiveSupport() on your SparkSession, Spark sets Hive as the default catalog for metadata lookups. Your SnappySession inherits this default catalog setting, which is why it’s pulling tables from Hive instead of SnappyData’s own metastore.

Solution 1: Use Separate Sessions for SnappyData and Hive

Create two independent SparkSession instances—one configured for SnappyData, another with Hive support. This keeps their metadata contexts isolated:

// Base SparkConf with SnappyData connection details
SparkConf sparkConf = new SparkConf()
    .setAppName("TEST APP")
    .set("snappydata.store.locators", "your-locator-host:10334") // Replace with your Snappy locator address
    // Add any other SnappyData-specific configs here

// Initialize SnappySession (uses SnappyData's metastore by default)
SnappySession snc = new SnappySession(
    SparkSession.builder()
        .config(sparkConf)
        .getOrCreate()
);

// Initialize Hive-enabled SparkSession (uses Hive metastore)
SparkSession hiveSession = SparkSession.builder()
    .config(sparkConf)
    .enableHiveSupport()
    .getOrCreate();

// Work with SnappyData tables
snc.sql("show tables").show(); // Lists SnappyData tables
Dataset<Row> snappyTableDF = snc.table("your_snappy_table");

// Work with Hive tables
hiveSession.sql("show tables").show(); // Lists Hive tables
Dataset<Row> hiveTableDF = hiveSession.table("your_hive_table");

Solution 2: Use Explicit Catalog References in a Single Session

SnappyData 1.0.1 supports multi-catalog functionality, so you can switch between SnappyData and Hive metastores in the same session by specifying the catalog name:

SparkConf sparkConf = new SparkConf()
    .setAppName("TEST APP")
    .set("snappydata.store.locators", "your-locator-host:10334")
    // Hive configs will be loaded automatically from hive-site.xml in classpath

// Create a single SparkSession with Hive support enabled
SparkSession spark = SparkSession.builder()
    .config(sparkConf)
    .enableHiveSupport()
    .getOrCreate();

SnappySession snc = new SnappySession(spark);

// Access SnappyData tables by explicitly using the "snappydata" catalog
snc.sql("USE CATALOG snappydata");
snc.sql("show tables").show(); // Lists SnappyData tables
Dataset<Row> snappyDF = snc.table("snappydata.default.your_snappy_table");

// Switch to Hive catalog to access Hive tables
snc.sql("USE CATALOG hive");
snc.sql("show tables").show(); // Lists Hive tables
Dataset<Row> hiveDF = snc.table("hive.default.your_hive_table");

Key Notes

  • Make sure hive-site.xml is in your application’s classpath when using Hive support.
  • Always specify your SnappyData locator address in the SparkConf—this ensures the SnappySession can connect to its cluster and metastore.
  • Avoid mixing multiple SparkContext instances; using separate SparkSessions (or catalog switching) is safer and more efficient.

内容的提问来源于stack exchange,提问作者umesh gowda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:19:28