You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为部署在OpenShift的Python Rest API添加存活探针以解决物化视图刷新卡顿问题

Absolutely! You can absolutely set up a liveness probe in OpenShift to handle this stuck refresh scenario—let’s break down how to do this properly, from the app side to the OpenShift config.

1. First, Build a Custom Health Check Endpoint in Your Python API

The key here is tracking the state of your materialized view refresh task, so the liveness probe can tell when it’s been stuck too long. Here’s a straightforward implementation using Flask (adjust for your framework like FastAPI if needed):

from flask import Flask, jsonify
from datetime import datetime
import time

app = Flask(__name__)

# Track refresh state (use Redis instead if running multi-process workers)
refresh_task = {
    "is_running": False,
    "start_timestamp": None
}

@app.route('/api/refresh-mv', methods=['POST'])
def trigger_refresh():
    global refresh_task
    refresh_task["is_running"] = True
    refresh_task["start_timestamp"] = datetime.now()
    
    try:
        # Your existing Oracle stored procedure call logic
        execute_oracle_refresh()
        # Reset state on success
        refresh_task["is_running"] = False
        refresh_task["start_timestamp"] = None
        return jsonify({"status": "refresh completed"}), 200
    except Exception as e:
        # Reset state on failure too
        refresh_task["is_running"] = False
        refresh_task["start_timestamp"] = None
        return jsonify({"error": str(e)}), 500

@app.route('/health/liveness', methods=['GET'])
def liveness_check():
    global refresh_task
    if refresh_task["is_running"]:
        elapsed_seconds = (datetime.now() - refresh_task["start_timestamp"]).total_seconds()
        # 2 hours = 7200 seconds
        if elapsed_seconds > 7200:
            # Return 5xx to signal the probe the container is unhealthy
            return jsonify({"status": "stuck", "elapsed": round(elapsed_seconds/60, 2)}), 500
    # All clear - return 200
    return jsonify({"status": "healthy"}), 200

def execute_oracle_refresh():
    # Replace with your actual stored procedure call
    # Example: cx_Oracle logic to call the refresh procedure
    pass

if __name__ == '__main__':
    app.run(host='0.0.0.0', port=8080)

Important note: If your app uses multi-process workers (like Gunicorn with multiple workers), the in-memory state won’t be shared across workers. Use a lightweight shared store like Redis to track the refresh state instead.

2. Configure the Liveness Probe in OpenShift

Now update your Deployment YAML to add the liveness probe that hits your new /health/liveness endpoint:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: mv-refresh-api
spec:
  replicas: 1
  selector:
    matchLabels:
      app: mv-refresh-api
  template:
    metadata:
      labels:
        app: mv-refresh-api
    spec:
      containers:
      - name: api-container
        image: your-python-api-image:latest
        ports:
        - containerPort: 8080
        livenessProbe:
          httpGet:
            path: /health/liveness
            port: 8080
          initialDelaySeconds: 30  # Wait 30s after container start before first check
          periodSeconds: 60       # Check every 60 seconds
          failureThreshold: 1     # Restart on first failed check
          timeoutSeconds: 5       # Fail check if no response in 5s

What these parameters do:

  • initialDelaySeconds: Gives your app time to fully start up before probe checks begin
  • periodSeconds: Balances between responsiveness and unnecessary overhead (60s is reasonable for a 2-hour threshold)
  • failureThreshold: Since we’re explicitly checking for a stuck state, one failure is enough to trigger a restart
  • timeoutSeconds: Prevents the probe itself from hanging if your app is unresponsive
3. Edge Cases & Best Practices
  • Graceful Shutdown: Add a SIGTERM handler in your Python app to clean up any partial work (if possible) before the container restarts
  • Logging: Add detailed logs for refresh start/end times and liveness probe results—this will help you debug why refreshes are getting stuck in the first place
  • Oracle Side Debugging: Don’t forget to check Oracle’s logs and performance metrics (like lock waits or long-running queries) to address the root cause of stuck refreshes, not just the symptom
  • Readiness Probe (Optional): If you want to stop routing traffic to the container while it’s refreshing, add a readiness probe that returns 5xx when the refresh is in progress

内容的提问来源于stack exchange,提问作者Ketan_Gupta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 11:18:16