You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS Elastic Beanstalk中Flask应用报HTTP 503且ELB健康异常求助

问题:AWS Elastic Beanstalk部署Flask Docker镜像后环境健康状态严重

ELB health is failing or not available for all instances.

部署仅含index.html的示例程序可正常运行,推测问题出在代码或镜像配置层面,相关代码及配置如下:

Flask应用代码:app.py

from flask import Flask, jsonify, request
from util import prediction

application = Flask(__name__)

@application.route('/predict', methods=['POST'])
def predict():
   data = request.get_json()
   try:
       sample = data['text']     
   except KeyError:
       return jsonify({'error':'No text sent'})
   pred = prediction(sample)
   try:
       result = jsonify(pred)
   except TypeError as e:
       result = jsonify({'error': str(e)})
   return result
   
if __name__ == '__main__':
   application.run(host='0.0.0.0', debug= True)

工具模块:util.py

import nltk
import pandas as pd
from nltk import TweetTokenizer
import numpy as np
import nltk
from nltk.stem.wordnet import WordNetLemmatizer
from sklearn.feature_extraction.text import TfidfVectorizer
import csv
import pandas as pd
import time
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.metrics import confusion_matrix, accuracy_score, classification_report
from nltk.tokenize import TweetTokenizer
from nltk.tag import pos_tag
import re
import string
from nltk.stem.wordnet import WordNetLemmatizer
from nltk.corpus import stopwords
import joblib
import warnings
warnings.filterwarnings("ignore")


# nltk.download('averaged_perceptron_tagger')
# nltk.download('wordnet')
# nltk.download('omw-1.4')
# nltk.download('stopwords')

token = TweetTokenizer()


def lemmatize_sentence(tokens):
    lemmatizer = WordNetLemmatizer()
    lemmatize_sentence = []
    for word, tag in pos_tag(tokens):
        if tag.startswith('NN'):
            pos = 'n'
        elif tag.startswith('VB'):
            pos = 'v'
        else:
            pos = 'a'
        lemmatize_sentence.append(lemmatizer.lemmatize(word, pos))
    return lemmatize_sentence
# print(' '.join(lemmatize_sentence(data[0][0])))


# Data cleaning, getting rid of words not needed for analysis.

stop_words = stopwords.words('english')


def cleaned(token):
    if token == 'u':
        return 'you'
    if token == 'r':
        return 'are'
    if token == 'some1':
        return 'someone'
    if token == 'yrs':
        return 'years'
    if token == 'hrs':
        return 'hours'
    if token == 'mins':
        return 'minutes'
    if token == 'secs':
        return 'seconds'
    if token == 'pls' or token == 'plz':
        return 'please'
    if token == '2morow':
        return 'tomorrow'
    if token == '2day':
        return 'today'
    if token == '4got' or token == '4gotten':
        return 'forget'
    if token == 'amp' or token == 'quot' or token == 'lt' or token == 'gt':
        return ''
    return token


# Noise removal from data, removing links, mentions and words with less than 3 length.


def remove_noise(tokens):
    cleaned_tokens = []
    for token, tag in pos_tag(tokens):
        # using non capturing groups ?:)// and eleminating the token if its a link.
        token = re.sub('http[s]?://(?:[a-zA-Z]|[0-9]|[$-_@.&+#]|[!*\(\),]|(?:%[0-9a-fA-F]))+', '', token)
        token = re.sub('[^a-zA-Z]', ' ', token)
        # eliminating token if its a mention
        token = re.sub("(@[A-Za-z0-9_]+)", "", token)
        if tag.startswith("NN"):
            pos = 'n'
        elif tag.startswith("VB"):
            pos = 'v'
        else:
            pos = 'a'

        lemmatizer = WordNetLemmatizer()
        token = lemmatizer.lemmatize(token, pos)

        cleaned_token = cleaned(token.lower())
        # Eliminating if the length of the token is less than 3, if its a punctuation or if it is a stopword.
        if cleaned_token not in string.punctuation and len(cleaned_token) > 2 and cleaned_token not in stop_words:
            cleaned_tokens.append(cleaned_token)
    return cleaned_tokens

with open ('Models/Sentimenttfpipe', 'rb') as f:
    loaded_pipeline = joblib.load(f)


def prediction(body):    
    # loaded_pipeline = joblib.load('Api/Models/Sentimenttfpipe')
    text= []
    test = token.tokenize(body)
    test = remove_noise(test)
    text.append(" ".join(test))
    test = pd.DataFrame(text, columns=['text'])
    a = loaded_pipeline.predict(test['text'].values.astype('U'))
    final = []
    if a[0] == 0:
        final.append({'Label' : 'Relaxed'})
        return {'Label' : 'Relaxed'}
        
    if a[0] == 1:
        final.append({'Label' : 'Angry'})
        return {'Label' : 'Angry'}
        
    if a[0] == 2:
        final.append({'Label' : 'Fearful'})
        return {'Label' : 'Fearful'}
        
    if a[0] == 3:
        final.append({'Label' : 'Happy'})
        return {'Label' : 'Happy'}
        
    if a[0] == 4:
        final.append({'Label' : 'Sad'})
        return {'Label' : 'Sad'}
        
    if a[0] == 5:
        final.append({'Label' : 'Surprised'})
        return {'Label' : 'Surprised'}
    
if __name__ == '__main__':
    sen = "May the force be with you"
    a = prediction(sen)
    print(a)

Dockerfile

FROM python:3.10.8

WORKDIR /app

COPY ["requirements.txt", "./"]

RUN pip install -r requirements.txt

RUN python -c "import nltk; nltk.download('averaged_perceptron_tagger'); nltk.download('wordnet'); nltk.download('omw-1.4'); nltk.download('stopwords');"

COPY . . 

EXPOSE 5000

ENTRYPOINT [ "gunicorn", "--bind=0.0.0.0:5000", "app:application" ]

docker-compose.yml

version: "3.7"

services:
  mlapp:
    container_name: Container
    image: mlapp
    ports: 
      - "5000:5000"
    build: 
      context: .
      dockerfile: Dockerfile

requirements.txt

Flask>=2.2.2
joblib==1.2.0
nltk==3.7
numpy==1.21.6
pandas==1.5.1
regex==2022.10.31
requests==2.28.1
scikit-learn==1.1.3
gunicorn==20.1.0

排查方向与解决方案

1. 健康检查端点缺失

AWS Elastic Beanstalk默认会对根路径/执行健康检查,但你的Flask应用只配置了/predict接口,没有处理GET请求的根路径。当ELB尝试访问/时,会返回404错误,导致健康状态判定为严重。

解决方法:添加一个根路径的GET接口用于健康检查:

@application.route('/')
def health_check():
    return jsonify({'status': 'healthy'}), 200

2. 模型文件路径问题

在util.py中,加载模型的路径是Models/Sentimenttfpipe,需要确认该目录和文件是否在Docker镜像中正确复制。如果构建镜像时Models目录未被包含,启动时会抛出文件找不到的异常,导致服务无法启动。

验证方法:在本地构建镜像后,进入容器检查文件是否存在:

docker run -it --rm mlapp ls /app/Models

如果路径错误或文件缺失,修正路径,确保Models目录被正确复制到镜像中(当前Dockerfile的COPY . .应该会包含,但需确认本地目录结构是否正确)。

3. 依赖安装与NLTK数据下载顺序

当前Dockerfile中先安装依赖再下载NLTK数据,这没问题,但需要确认下载是否成功。可以在构建镜像时添加日志输出,或进入容器验证NLTK数据是否存在:

docker run -it --rm mlapp python -c "import nltk; print(nltk.data.find('corpora/stopwords'))"

如果下载失败,可尝试更换NLTK数据源镜像,或在下载命令中指定--timeout参数:

RUN python -c "import nltk; nltk.download('averaged_perceptron_tagger', timeout=60); nltk.download('wordnet', timeout=60); nltk.download('omw-1.4', timeout=60); nltk.download('stopwords', timeout=60);"

4. Gunicorn启动参数与端口配置

确认Gunicorn绑定的端口是5000,且Elastic Beanstalk的容器端口配置与之一致(默认情况下,Docker部署会自动映射EXPOSE的端口,但可以在EB配置中明确指定容器端口为5000)。

另外,可在Gunicorn启动命令中添加--workers参数,提升服务稳定性:

ENTRYPOINT [ "gunicorn", "--bind=0.0.0.0:5000", "--workers=2", "app:application" ]

5. 查看Elastic Beanstalk日志

直接查看EB环境的日志是最有效的排查方式,可通过EB控制台下载完整日志,重点查看:

  • 容器启动日志:确认服务是否正常启动
  • 应用日志:查看是否有Python异常抛出

内容的提问来源于stack exchange,提问作者M Bilal Ayaz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 04:20:23