AWS Fargate环境下Spark无法解析本地IP问题求助
Let's break down your issue first: the error java.net.UnknownHostException: d4771b650361: d4771b650361: Name or service not known tells us that Spark (via Java) can't resolve your container's random hostname (the auto-generated container ID) to an IP address. This is a common gotcha with Fargate because by default, the container's /etc/hosts file doesn't map this auto-generated hostname to a valid IP.
Here are targeted fixes to resolve this:
1. Force Spark to Use a Known IP Instead of Hostname
Spark relies on Java's hostname resolution during initialization. Bypass this by explicitly telling Spark and Java to use your ECS metadata IP:
- Update your
$SPARK_HOME/conf/spark-env.shwith these lines (replace<ECS_METADATA_IP>with the IP you fetch fromhttp://169.254.170.2/v2/metadata):export SPARK_LOCAL_IP=<ECS_METADATA_IP> export SPARK_MASTER_HOST=<ECS_METADATA_IP> export JAVA_OPTS="-Djava.net.preferIPv4Stack=true -Djava.net.hostname=<ECS_METADATA_IP>" - Alternatively, pass these configs directly in your
spark-submitcommand to override defaults:spark-submit \ --master spark://<ECS_METADATA_IP>:7077 \ --conf spark.driver.host=<ECS_METADATA_IP> \ --conf spark.driver.bindAddress=<ECS_METADATA_IP> \ --verbose \ --jars lib/RedshiftJDBC42-1.2.12.1017.jar \ --packages org.apache.hadoop:hadoop-aws:2.7.3,com.amazonaws:aws-java-sdk:1.7.4,com.upplication:s3fs:2.2.1 \ ./build_phase.py
2. Fix /etc/hosts to Map Container Hostname to Valid IP
Your current /etc/hosts update only handles the master mapping—you also need to map the container's own hostname to an IP. Add this to your setup script:
# Get the container's auto-generated hostname (usually the container ID) CONTAINER_HOSTNAME=$(hostname) # Fetch your ECS task's private IP from metadata ECS_IP=$(curl -s http://169.254.170.2/v2/metadata | awk -F'"' '/IPv4Addresses/{print $4}') # Update /etc/hosts to resolve the container hostname echo "$ECS_IP $CONTAINER_HOSTNAME" >> /etc/hosts echo "127.0.0.1 localhost $CONTAINER_HOSTNAME" >> /etc/hosts
Note: If your image doesn't have awk, you can use grep or install jq for cleaner JSON parsing.
3. Optimize for Single-Node Fargate with local[*] Mode
Since you're running on a single Fargate node, using Spark's local mode is often simpler than setting up a standalone master. Just ensure you explicitly set the driver host to avoid hostname resolution issues:
spark-submit \ --master local[*] \ --conf spark.driver.host=<ECS_METADATA_IP> \ --conf spark.driver.bindAddress=0.0.0.0 \ --jars lib/RedshiftJDBC42-1.2.12.1017.jar \ --packages org.apache.hadoop:hadoop-aws:2.7.3,com.amazonaws:aws-java-sdk:1.7.4,com.upplication:s3fs:2.2.1 \ ./build_phase.py
4. Verify Fargate Network Config
Even though you mentioned your subnet has a NAT Gateway, double-check:
- Your task definition's security group allows inbound/outbound traffic on Spark ports (7077, 4040, etc.) for the container itself.
- If you set a custom
hostnamein your Fargate task definition, make sure/etc/hostsmaps that name to your ECS IP.
The root cause here is that Fargate assigns a random container ID as the hostname by default, and Java's InetAddress.getLocalHost() (called during Spark's initialization) fails to resolve it because there's no entry in /etc/hosts. The fixes above either bypass hostname resolution entirely or fix the mapping so Java can resolve it correctly.
内容的提问来源于stack exchange,提问作者Adam

