ejabberd升级后oom_killer误杀进程问题求助
Let's break down your issue and walk through actionable fixes based on the details you provided:
Problem Recap
After upgrading from ejabberd 16 to 18.04, you're seeing processes get killed with the message:
Killed 1 process(es) consuming more than 38108 message(s) each
This happens even with 7GB of free RAM on your 12GB Ubuntu 16.04 VM. Disabling oom_killer stops the kills but leads to excessive swap usage, and you've confirmed versions 17.12+ (including 18.04) have this issue, while 17.01 and earlier work fine. The core architectural bottleneck is a high-traffic chat room where a Python bot floods thousands of messages per minute, and the archiving process can't keep up with the load.
Why This Is Happening
The key change here is ejabberd 17.12's updated oom_killer — it doesn't trigger based on system memory usage, but on per-process message queue length. The default threshold (38108 messages) is being hit because your archiving process can't process the flood of messages from the aggregated chat room, causing queues to balloon. The threshold parameter exists in the source code but isn't documented, which is why you couldn't find it in official docs.
Actionable Solutions
1. Tweak the Undocumented oom_killer Threshold
Even though it's not officially documented, you can adjust the message queue threshold to better match your system's processing capacity:
- Open your ejabberd configuration file (usually
/etc/ejabberd/ejabberd.yml) - Add or update the
oom_killer_thresholdparameter to a higher value (e.g., 100000) alongside enablingoom_killer:oom_killer: true oom_killer_threshold: 100000 - Test this value gradually — set it high enough to avoid false kills but low enough to prevent unbounded memory growth. Note that since this is an undocumented parameter, it might change in future ejabberd versions, so keep an eye on release notes.
2. Fix the Root Architectural Bottleneck
The oom_killer is just a symptom of message queue overflow. Addressing the source will prevent the issue long-term:
- Remove the aggregated chat room: Split the data into multiple smaller rooms or topics. This distributes the message processing load across multiple processes instead of overwhelming a single one.
- Optimize the archiving process:
- Increase the number of archiving processes using ejabberd's
max_fsm_queueparameter to boost concurrent message handling. - Add rate limiting to your Python bot: Throttle message sends to match the archiving process's capacity (e.g., limit to a few hundred messages per minute instead of thousands).
- Increase the number of archiving processes using ejabberd's
- Check for lingering memory leaks: While 18.01 fixed 17.12's known leak, monitor ejabberd's memory usage over time with
top,htop, orejabberdctl statusto rule out other leaks.
3. Environment & Runtime Tuning
- Test alternative OS environments: As you planned, try FreeBSD and Ubuntu 18.04. Different OS kernels and Erlang runtime environments can handle message queue pressure differently, so you might see better stability elsewhere.
- Optimize Erlang VM parameters: Tune the Erlang VM to handle higher message throughput:
- Set
+P 100000to increase the maximum number of processes. - Enable asynchronous IO with
+K trueand adjust async thread count with+A 32(adjust based on your CPU cores). - These parameters can be added to the ejabberd startup script or systemd service file.
- Set
- Add monitoring: Track message queue lengths, process memory, and swap usage to catch bottlenecks early. Tools like
ejabberdctl mnesia statuscan give you insight into internal queues, while system-level tools likevmstatoriostathelp monitor resource usage.
内容的提问来源于stack exchange,提问作者Dreddnar

