AWS Elastic Beanstalk中Flask应用报HTTP 503且ELB健康异常求助
ELB health is failing or not available for all instances.
部署仅含index.html的示例程序可正常运行,推测问题出在代码或镜像配置层面,相关代码及配置如下:
Flask应用代码:app.py
from flask import Flask, jsonify, request from util import prediction application = Flask(__name__) @application.route('/predict', methods=['POST']) def predict(): data = request.get_json() try: sample = data['text'] except KeyError: return jsonify({'error':'No text sent'}) pred = prediction(sample) try: result = jsonify(pred) except TypeError as e: result = jsonify({'error': str(e)}) return result if __name__ == '__main__': application.run(host='0.0.0.0', debug= True)
工具模块:util.py
import nltk import pandas as pd from nltk import TweetTokenizer import numpy as np import nltk from nltk.stem.wordnet import WordNetLemmatizer from sklearn.feature_extraction.text import TfidfVectorizer import csv import pandas as pd import time from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.linear_model import LogisticRegression import pandas as pd from sklearn.model_selection import train_test_split from sklearn.metrics import confusion_matrix, accuracy_score, classification_report from nltk.tokenize import TweetTokenizer from nltk.tag import pos_tag import re import string from nltk.stem.wordnet import WordNetLemmatizer from nltk.corpus import stopwords import joblib import warnings warnings.filterwarnings("ignore") # nltk.download('averaged_perceptron_tagger') # nltk.download('wordnet') # nltk.download('omw-1.4') # nltk.download('stopwords') token = TweetTokenizer() def lemmatize_sentence(tokens): lemmatizer = WordNetLemmatizer() lemmatize_sentence = [] for word, tag in pos_tag(tokens): if tag.startswith('NN'): pos = 'n' elif tag.startswith('VB'): pos = 'v' else: pos = 'a' lemmatize_sentence.append(lemmatizer.lemmatize(word, pos)) return lemmatize_sentence # print(' '.join(lemmatize_sentence(data[0][0]))) # Data cleaning, getting rid of words not needed for analysis. stop_words = stopwords.words('english') def cleaned(token): if token == 'u': return 'you' if token == 'r': return 'are' if token == 'some1': return 'someone' if token == 'yrs': return 'years' if token == 'hrs': return 'hours' if token == 'mins': return 'minutes' if token == 'secs': return 'seconds' if token == 'pls' or token == 'plz': return 'please' if token == '2morow': return 'tomorrow' if token == '2day': return 'today' if token == '4got' or token == '4gotten': return 'forget' if token == 'amp' or token == 'quot' or token == 'lt' or token == 'gt': return '' return token # Noise removal from data, removing links, mentions and words with less than 3 length. def remove_noise(tokens): cleaned_tokens = [] for token, tag in pos_tag(tokens): # using non capturing groups ?:)// and eleminating the token if its a link. token = re.sub('http[s]?://(?:[a-zA-Z]|[0-9]|[$-_@.&+#]|[!*\(\),]|(?:%[0-9a-fA-F]))+', '', token) token = re.sub('[^a-zA-Z]', ' ', token) # eliminating token if its a mention token = re.sub("(@[A-Za-z0-9_]+)", "", token) if tag.startswith("NN"): pos = 'n' elif tag.startswith("VB"): pos = 'v' else: pos = 'a' lemmatizer = WordNetLemmatizer() token = lemmatizer.lemmatize(token, pos) cleaned_token = cleaned(token.lower()) # Eliminating if the length of the token is less than 3, if its a punctuation or if it is a stopword. if cleaned_token not in string.punctuation and len(cleaned_token) > 2 and cleaned_token not in stop_words: cleaned_tokens.append(cleaned_token) return cleaned_tokens with open ('Models/Sentimenttfpipe', 'rb') as f: loaded_pipeline = joblib.load(f) def prediction(body): # loaded_pipeline = joblib.load('Api/Models/Sentimenttfpipe') text= [] test = token.tokenize(body) test = remove_noise(test) text.append(" ".join(test)) test = pd.DataFrame(text, columns=['text']) a = loaded_pipeline.predict(test['text'].values.astype('U')) final = [] if a[0] == 0: final.append({'Label' : 'Relaxed'}) return {'Label' : 'Relaxed'} if a[0] == 1: final.append({'Label' : 'Angry'}) return {'Label' : 'Angry'} if a[0] == 2: final.append({'Label' : 'Fearful'}) return {'Label' : 'Fearful'} if a[0] == 3: final.append({'Label' : 'Happy'}) return {'Label' : 'Happy'} if a[0] == 4: final.append({'Label' : 'Sad'}) return {'Label' : 'Sad'} if a[0] == 5: final.append({'Label' : 'Surprised'}) return {'Label' : 'Surprised'} if __name__ == '__main__': sen = "May the force be with you" a = prediction(sen) print(a)
Dockerfile
FROM python:3.10.8 WORKDIR /app COPY ["requirements.txt", "./"] RUN pip install -r requirements.txt RUN python -c "import nltk; nltk.download('averaged_perceptron_tagger'); nltk.download('wordnet'); nltk.download('omw-1.4'); nltk.download('stopwords');" COPY . . EXPOSE 5000 ENTRYPOINT [ "gunicorn", "--bind=0.0.0.0:5000", "app:application" ]
docker-compose.yml
version: "3.7" services: mlapp: container_name: Container image: mlapp ports: - "5000:5000" build: context: . dockerfile: Dockerfile
requirements.txt
Flask>=2.2.2 joblib==1.2.0 nltk==3.7 numpy==1.21.6 pandas==1.5.1 regex==2022.10.31 requests==2.28.1 scikit-learn==1.1.3 gunicorn==20.1.0
排查方向与解决方案
1. 健康检查端点缺失
AWS Elastic Beanstalk默认会对根路径/执行健康检查,但你的Flask应用只配置了/predict接口,没有处理GET请求的根路径。当ELB尝试访问/时,会返回404错误,导致健康状态判定为严重。
解决方法:添加一个根路径的GET接口用于健康检查:
@application.route('/') def health_check(): return jsonify({'status': 'healthy'}), 200
2. 模型文件路径问题
在util.py中,加载模型的路径是Models/Sentimenttfpipe,需要确认该目录和文件是否在Docker镜像中正确复制。如果构建镜像时Models目录未被包含,启动时会抛出文件找不到的异常,导致服务无法启动。
验证方法:在本地构建镜像后,进入容器检查文件是否存在:
docker run -it --rm mlapp ls /app/Models
如果路径错误或文件缺失,修正路径,确保Models目录被正确复制到镜像中(当前Dockerfile的COPY . .应该会包含,但需确认本地目录结构是否正确)。
3. 依赖安装与NLTK数据下载顺序
当前Dockerfile中先安装依赖再下载NLTK数据,这没问题,但需要确认下载是否成功。可以在构建镜像时添加日志输出,或进入容器验证NLTK数据是否存在:
docker run -it --rm mlapp python -c "import nltk; print(nltk.data.find('corpora/stopwords'))"
如果下载失败,可尝试更换NLTK数据源镜像,或在下载命令中指定--timeout参数:
RUN python -c "import nltk; nltk.download('averaged_perceptron_tagger', timeout=60); nltk.download('wordnet', timeout=60); nltk.download('omw-1.4', timeout=60); nltk.download('stopwords', timeout=60);"
4. Gunicorn启动参数与端口配置
确认Gunicorn绑定的端口是5000,且Elastic Beanstalk的容器端口配置与之一致(默认情况下,Docker部署会自动映射EXPOSE的端口,但可以在EB配置中明确指定容器端口为5000)。
另外,可在Gunicorn启动命令中添加--workers参数,提升服务稳定性:
ENTRYPOINT [ "gunicorn", "--bind=0.0.0.0:5000", "--workers=2", "app:application" ]
5. 查看Elastic Beanstalk日志
直接查看EB环境的日志是最有效的排查方式,可通过EB控制台下载完整日志,重点查看:
- 容器启动日志:确认服务是否正常启动
- 应用日志:查看是否有Python异常抛出
内容的提问来源于stack exchange,提问作者M Bilal Ayaz

