You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

高负载下OpenvSwitch端口丢失问题排查咨询

Hey there, let's dive into your OVS + OpenStack HA VIP issue step by step:

复现过程中及后续的必要操作

To make sure you capture all critical data when the issue reproduces, here's what you should do:

  • Before reproducing

    • Back up your current OVS configuration first: run ovs-vsctl show > ovs-config-pre-repro.txt to save the full port/bridge state.
    • Archive existing OVS logs (e.g., /var/log/openvswitch/ovs-vswitchd.log) to a separate file—debug logs will generate a ton of data, and you don't want to overwrite the original warning logs you already found.
    • Note down the exact name of your VIP-associated OVS port, so you can filter logs quickly later.
  • During reproduction

    • Continuously monitor the VIP port's existence with a script or repeated command: for example, watch -n 2 "ovs-vsctl list port | grep -A 5 -B 5 'your-vip-port-name'" to track if it disappears.
    • Keep an eye on system resource usage: use top/htop to check ovs-vswitchd's CPU/memory consumption, and vmstat/iostat to monitor overall system load and IO performance—high resource contention is often a trigger here.
    • Filter the debug logs in real-time to avoid being overwhelmed: run tail -f /var/log/openvswitch/ovs-vswitchd.log | grep -E "(your-vip-port-name|revalidator|port.*deleted|Unreasonably long)" to focus on relevant entries.
  • After the issue occurs

    • Immediately freeze the system state: stop all ongoing OpenStack operations to prevent further state changes.
    • Export the post-failure OVS configuration: ovs-vsctl show > ovs-config-post-failure.txt to compare with the pre-repro state.
    • Collect all debug logs, plus system logs (e.g., /var/log/messages or /var/log/syslog)—look for entries around the time the port disappeared that mention OVS threads, lock contention, or resource exhaustion.
    • Use ovs-appctl vlog/save ovs-debug-log-snapshot.txt to save a snapshot of the current debug logs, in case the log file gets rotated or overwritten.
Is this issue typical? What causes it?

This is indeed a known scenario with OVS in high-load environments, especially when dealing with HA VIPs managed via OVS ports. The warning you saw—Unreasonably long 62741ms poll interval—is the key clue here.

The revalidator thread in ovs-vswitchd is responsible for syncing the internal OVS state with the actual system/network state. When this thread gets stuck or delayed for an abnormally long time, several bad things can happen:

  1. Thread scheduling blockage: High concurrency (like your volume-from-glance-image operations) generates a flood of OVS configuration changes or network events. The revalidator thread might get stuck waiting for locks, or starved of CPU time by other processes, leading to the extended poll interval.
  2. State inconsistency: When the revalidator can't keep up, OVS might misinterpret the VIP port as "unhealthy" or "non-existent" and trigger its cleanup logic, deleting the port entirely.
  3. Legacy OVS bugs: Older OVS versions (pre-2.8.x, especially 2.7 and earlier) had known issues with revalidator thread handling under high load, including incorrect port state checks that led to accidental port deletion.
About your search direction question

Your observation is spot-on: most public OVS content focuses on data plane issues (flow tables, packet forwarding, etc.), but your problem lies in the control plane—specifically the internal thread management and state maintenance in ovs-vswitchd.

To find relevant info, adjust your search keywords to target this layer: try terms like ovs revalidator thread port missing, ovs-vswitchd port disappear high load, or Unreasonably long poll interval port deleted. You'll likely find community discussions, bug reports, and even patch notes that address this exact scenario.


内容的提问来源于stack exchange,提问作者Wey Gu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:33:52