You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark Shell集成spark-cassandra-connector 3.1.0启动失败求助

Troubleshooting Spark-Cassandra Connector Socket Failures & Working Setup Guide

Let's break down why you're hitting that ClosedChannelException and walk through how to get your Spark Shell with Cassandra connector up and running properly.

Why This Is Happening

  1. Unreachable Cassandra Node IP
    The IP 172.20.12.135 in your error is a private/internal address from your Cassandra cluster, but your Spark client machine can't resolve or reach it (hence the nslookup failure). This usually happens if Cassandra is configured to broadcast its internal private IP instead of a public/resolvable address that Spark can access.
  2. Port or Network Blockage
    The non-default port 64320 might be blocked by firewalls, security groups, or network policies. Even if the connection starts, it gets closed immediately, triggering the ClosedChannelException.
  3. Version Check (Good News!)
    Your setup uses Spark 3.1.2 with connector 3.1.0—this is fully compatible (connector 3.x matches Spark 3.x), so version mismatch isn't the issue here.

Quick Fix: Force Spark to Use Accessible Cassandra Addresses

Instead of relying on automatic node discovery (which is pulling that unresolvable IP), explicitly tell Spark where to find your Cassandra cluster when launching the shell.

Option 1: Using --packages

spark-shell --packages com.datastax.spark:spark-cassandra-connector_2.12:3.1.0 \
  --conf spark.cassandra.connection.host=YOUR_CASSANDRA_PUBLIC_IP_OR_HOSTNAME \
  --conf spark.cassandra.connection.port=YOUR_CASSANDRA_NATIVE_PORT # Default is 9042—use your actual port if changed

Option 2: Using Local Assembly Jar

spark-shell --jars ~/.asdf/installs/spark/3.1.2/jars/spark-cassandra-connector-assembly-3.1.0-11-g53f24ce9.jar \
  --conf spark.cassandra.connection.host=YOUR_CASSANDRA_PUBLIC_IP_OR_HOSTNAME \
  --conf spark.cassandra.connection.port=YOUR_CASSANDRA_NATIVE_PORT

Permanent Fix: Fix Cassandra's Broadcast Configuration

If you want Spark to automatically discover all Cassandra nodes without manual IP entries, update your Cassandra cluster's configuration:

  1. On each Cassandra node, open the cassandra.yaml file (typically in /etc/cassandra/ or your installation's conf/ directory).
  2. Adjust these settings:
    • Set rpc_address to 0.0.0.0 (allows the node to accept RPC requests from any address)
    • Set broadcast_rpc_address to the public IP or resolvable hostname of the node (not the private internal IP like 172.20.12.135)
    • Double-check native_transport_port (default 9042) and ensure this port is open in firewalls/security groups for Spark's network
  3. Restart all Cassandra nodes to apply the changes.

Verify the Connection

Once you launch the Spark Shell with the corrected settings, test the connector with a simple read operation:

// Replace "your_keyspace" and "your_table" with your actual Cassandra resources
val testDF = spark.read.format("org.apache.spark.sql.cassandra")
  .options(Map("keyspace" -> "your_keyspace", "table" -> "your_table"))
  .load()

// Print the first 5 rows to confirm connectivity
testDF.show(5)

If this runs without errors and shows your data, you're all set!

内容的提问来源于stack exchange,提问作者Martin Macak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 15:02:37