在Heroku部署RASA聊天机器人遭遇R14/R10错误求助
Heroku部署Rasa聊天机器人遇到的模型加载问题
问题背景
我正在Heroku上部署RASA聊天机器人,用了下面的Dockerfile和脚本文件。服务器看着像是正常启动了,但我感觉模型提取环节肯定出问题了,目前搞不定这个状况。
用到的配置文件
Dockerfile
FROM rasa/rasa:1.9.7 USER root ENTRYPOINT [] ADD . /app/ RUN pip install rasa[spacy] RUN python -m spacy download en_core_web_md RUN python -m spacy link en_core_web_md en RUN ls /app RUN chmod +x /app/server.sh CMD /app/server.sh credentials.yml --debug
server.sh脚本
rasa run -m models --cors "*" --enable-api --port $PORT --endpoints endpoints.yml --credentials credentials.yml --debug
Heroku实时日志(heroku logs -t)
2020-06-09T11:25:15.037081+00:00 heroku[web.1]: State changed from crashed to starting 2020-06-09T11:25:49.798297+00:00 heroku[web.1]: Starting process with command `/bin/bash -o pipefail -c /app/server.sh` 2020-06-09T11:25:56.027748+00:00 app[web.1]: 2020-06-09 11:25:56 DEBUG rasa.core.utils - Available web server routes: 2020-06-09T11:25:56.027812+00:00 app[web.1]: /conversations/<conversation_id>/messages POST add_message 2020-06-09T11:25:56.027813+00:00 app[web.1]: /conversations/<conversation_id>/tracker/events POST append_events 2020-06-09T11:25:56.027815+00:00 app[web.1]: /webhooks/rasa GET custom_webhook_RasaChatInput.health 2020-06-09T11:25:56.027815+00:00 app[web.1]: /webhooks/rasa/webhook POST custom_webhook_RasaChatInput.receive 2020-06-09T11:25:56.027816+00:00 app[web.1]: /webhooks/rest GET custom_webhook_RestInput.health 2020-06-09T11:25:56.027816+00:00 app[web.1]: /webhooks/rest/webhook POST custom_webhook_RestInput.receive 2020-06-09T11:25:56.027824+00:00 app[web.1]: /model/test/intents POST evaluate_intents 2020-06-09T11:25:56.027825+00:00 app[web.1]: /model/test/stories POST evaluate_stories 2020-06-09T11:25:56.027825+00:00 app[web.1]: /conversations/<conversation_id>/execute POST execute_action 2020-06-09T11:25:56.027826+00:00 app[web.1]: /domain GET get_domain 2020-06-09T11:25:56.027826+00:00 app[web.1]: /socket.io POST handle_request 2020-06-09T11:25:56.027826+00:00 app[web.1]: / GET hello 2020-06-09T11:25:56.027827+00:00 app[web.1]: /model PUT load_model 2020-06-09T11:25:56.027827+00:00 app[web.1]: /model/parse POST parse 2020-06-09T11:25:56.027828+00:00 app[web.1]: /conversations/<conversation_id>/predict POST predict 2020-06-09T11:25:56.027829+00:00 app[web.1]: /conversations/<conversation_id>/tracker/events PUT replace_events 2020-06-09T11:25:56.027830+00:00 app[web.1]: /conversations/<conversation_id>/story GET retrieve_story 2020-06-09T11:25:56.027830+00:00 app[web.1]: /conversations/<conversation_id>/tracker GET retrieve_tracker 2020-06-09T11:25:56.027830+00:00 app[web.1]: /webhooks/socketio GET socketio_webhook.health 2020-06-09T11:25:56.027831+00:00 app[web.1]: /status GET status 2020-06-09T11:25:56.027831+00:00 app[web.1]: /model/predict POST tracker_predict 2020-06-09T11:25:56.027832+00:00 app[web.1]: /model/train POST train 2020-06-09T11:25:56.027832+00:00 app[web.1]: /conversations/<conversation_id>/trigger_intent POST trigger_intent 2020-06-09T11:25:56.027832+00:00 app[web.1]: /model DELETE unload_model 2020-06-09T11:25:56.027833+00:00 app[web.1]: /version GET version 2020-06-09T11:25:56.027882+00:00 app[web.1]: 2020-06-09 11:25:56 INFO root - Starting Rasa server on http://localhost:17793 2020-06-09T11:25:56.028133+00:00 app[web.1]: 2020-06-09 11:25:56 DEBUG rasa.core.utils - Using the default number of Sanic workers (1). 2020-06-09T11:25:56.029616+00:00 app[web.1]: 2020-06-09 11:25:56 INFO root - Enabling coroutine debugging. Loop id 94472838346280. 2020-06-09T11:25:56.345896+00:00 app[web.1]: 2020-06-09 11:25:56 DEBUG rasa.model - Extracted model to '/tmp/tmp6ebr6r7s'. 2020-06-09T11:26:03.846339+00:00 heroku[web.1]: Process running mem=685M(133.9%) 2020-06-09T11:26:03.848858+00:00 heroku[web.1]: Error R14(内存配额超限) 2020-06-09T11:26:25.775231+00:00 heroku[web.1]: Process running mem=909M(177.7%) 2020-06-09T11:26:25.777768+00:00 heroku[web.1]: Error R14(内存配额超限) 2020-06-09T11:26:34.617203+00:00 app[web.1]: 2020-06-09 11:26:34 INFO rasa.nlu.components - Added 'SpacyNLP' to component cache. Key 'SpacyNLP-en'. 2020-06-09T11:26:34.715295+00:00 app[web.1]: 2020-06-09 11:26:34 DEBUG rasa.utils.tensorflow.models - Loading the model ... 2020-06-09T11:26:35.297208+00:00 app[web.1]: 2020-06-09 11:26:35.296743: E tensorflow/stream_executor/cuda/cuda_driver.cc:351] failed call to cuInit: UNKNOWN ERROR (303) 2020-06-09T11:26:41.647385+00:00 app[web.1]: 2020-06-09 11:26:41 DEBUG rasa.utils.tensorflow.models - Finished loading the model. 2020-06-09T11:26:41.648183+00:00 app[web.1]: 2020-06-09 11:26:41 DEBUG rasa.utils.tensorflow.models - Building tensorflow prediction graph... 2020-06-09T11:26:47.461993+00:00 heroku[web.1]: Process running mem=979M(191.2%) 2020-06-09T11:26:47.463550+00:00 heroku[web.1]: Error R14(内存配额超限) 2020-06-09T11:26:49.828941+00:00 heroku[web.1]: Error R10(启动超时)-> Web进程未在启动后60秒内绑定到$PORT的内容
问题分析与解决建议
从日志里能揪出几个核心问题:
- 内存超限(R14错误):模型加载和TensorFlow图构建的时候,内存占用直接飙到979M,远远超过Heroku默认的512M配额,这会导致进程被限制,严重的话直接被终止。
- 启动超时(R10错误):Heroku要求Web进程必须在60秒内绑定到指定端口,但模型加载和图构建耗时太长,直接超了时限。
- CUDA错误:日志里的
cuInit: UNKNOWN ERROR (303)其实不用管,因为Heroku的dyno没有GPU,TensorFlow会自动切换到CPU运行,属于正常情况。
针对这些问题,给你几个实用的解决方向:
- 优化内存占用:
- 换成更轻量的Spacy模型,比如用
en_core_web_sm代替en_core_web_md,能省不少内存。 - 在Rasa的配置文件里调整TensorFlow参数,比如把
batch_size设小一点,或者禁用一些用不上的组件。 - 提前在本地训练好模型,把
models目录直接打包到镜像里,避免容器启动时额外加载依赖导致内存暴涨。
- 换成更轻量的Spacy模型,比如用
- 解决启动超时:
- 升级Heroku的dyno配置,比如换成Standard 1X或2X,更大的内存和CPU能加快模型加载速度,避免超时。
- 把模型提取和加载的步骤放到镜像构建阶段,而不是容器启动时做。比如在Dockerfile里添加加载模型的命令,启动时直接运行服务器就行。
- 调整启动命令:
- 检查Dockerfile的CMD参数,你写的
CMD /app/server.sh credentials.yml --debug多传了参数,server.sh里已经指定了credentials.yml,改成CMD /app/server.sh --debug就行,避免参数冲突。
- 检查Dockerfile的CMD参数,你写的
内容的提问来源于stack exchange,提问作者sheel
相关产品推荐
相关产品推荐

