You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

异构配置Hadoop集群中Spark应用资源最优配置咨询

Hey there! Let's break down how to tune your Spark parameters for this heterogeneous cluster and fix that unexpected slowdown you're seeing.

First: Fix the Mode Conflict in Your Code

Your current code mixes two Spark cluster modes (spark://master:7077 for Standalone and yarn-cluster for YARN), which creates resource scheduling chaos and is almost certainly contributing to the slow performance. Stick with YARN mode—it’s far better at managing resources across heterogeneous nodes.

Here’s the cleaned-up code to start with:

from pyspark.sql import SparkSession

# Initialize SparkSession with YARN cluster mode (no duplicate contexts needed!)
spark = SparkSession.builder \
    .appName("SimpleApplication") \
    .master("yarn") \
    .config("spark.submit.deployMode", "cluster") \
    # We'll add resource configs here next
    .getOrCreate()

# You don't need separate SparkContext/HiveContext/SQLContext—SparkSession includes them
sc = spark.sparkContext
hc = spark._wrapped
sqlContext = spark._wrapped

Second: Calculate Usable Cluster Resources

First, we need to reserve resources for Hadoop’s core services (so they don’t compete with Spark):

  • Machine 1 (Master + Slave):
    • Total: 16GB RAM, 4 vCPUs
    • Reserve 2GB RAM + 1 vCPU for Master services (NameNode, ResourceManager)
    • Usable for Spark: 14GB RAM, 3 vCPUs
  • Machine 2 (Slave only):
    • Total: 8GB RAM, 2 vCPUs
    • Reserve 1GB RAM for Slave services (DataNode, NodeManager)
    • Usable for Spark: 7GB RAM, 2 vCPUs

Third: Tune Spark Resource Parameters

We’ll configure Spark to use every bit of usable resource while avoiding overloading nodes:

YARN Mode (Recommended for Heterogeneous Clusters)

YARN automatically adapts to node resource differences, so we can set parameters that align with our calculated usable resources:

spark = SparkSession.builder \
    .appName("SimpleApplication") \
    .master("yarn") \
    .config("spark.submit.deployMode", "cluster") \
    # 2 executors (one per node)
    .config("spark.executor.instances", "2") \
    # Let YARN assign cores based on node availability (3 on Machine1, 2 on Machine2)
    .config("spark.executor.cores", "3") \
    # YARN will adjust memory per executor: ~12GB on Machine1, ~6GB on Machine2
    .config("spark.executor.memory", "12g") \
    # Add overhead to prevent OOM (10% of executor memory is a safe bet)
    .config("spark.executor.memoryOverhead", "1g") \
    # Give Driver enough memory (runs on a YARN node in cluster mode)
    .config("spark.driver.memory", "2g") \
    # Set parallelism to 2-3x total cores (3+2=5 → 12 is a good middle ground)
    .config("spark.default.parallelism", "12") \
    # Match shuffle partitions to parallelism for SQL jobs
    .config("spark.sql.shuffle.partitions", "12") \
    .getOrCreate()

If you want even more control, you can configure YARN’s node-specific resource quotas (edit yarn-site.xml and capacity-scheduler.xml):

  • Set Machine1’s available resources to 14GB RAM / 3 vCPUs
  • Set Machine2’s available resources to 7GB RAM / 2 vCPUs
    Then Spark will automatically allocate the exact usable resources to each executor without extra tuning.

Standalone Mode (Not Recommended, But If You Prefer It)

For Standalone mode, first update each worker’s configs:

  • Machine1 Worker: Set SPARK_WORKER_CORES=3 and SPARK_WORKER_MEMORY=14g in spark-env.sh
  • Machine2 Worker: Set SPARK_WORKER_CORES=2 and SPARK_WORKER_MEMORY=7g
    Then use this Spark config:
spark = SparkSession.builder \
    .appName("SimpleApplication") \
    .master("spark://master:7077") \
    .config("spark.executor.instances", "2") \
    .config("spark.executor.cores", "3") \
    .config("spark.executor.memory", "12g") \
    .config("spark.default.parallelism", "12") \
    .getOrCreate()

Why Your Cluster Was Slower Than Single Node

  1. Mode Conflict: Mixing Standalone and YARN modes made Spark unable to properly schedule resources, so it was likely underutilizing your cluster.
  2. Unoptimized Parallelism: Default Spark settings use low parallelism, which means your cores were sitting idle while waiting for tasks.
  3. Resource Competition: Without reserving resources for Hadoop services, your Master node’s core services were fighting with Spark for CPU/RAM, slowing everything down.

内容的提问来源于stack exchange,提问作者Arun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:46:02