关于在CloudSim中无优化算法实现重试FTT及算法修改的技术咨询
Hey there! Let's tackle your two CloudSim-related questions head-on— I’ve spent a fair amount of time working with its fault tolerance features, so I can walk you through this.
A straightforward retry strategy doesn’t require complex optimization algorithms— you just need to hook into CloudSim’s event system to detect failures and resubmit cloudlets. Here’s how to do it:
Step 1: Track Failure Events
CloudSim uses an event-driven model, so you’ll want to extend theDatacenterBrokerclass to monitor cloudlet completion statuses. Focus on catchingFAILEDorFAILED_RESOURCE_UNAVAILABLEstatuses in theonCloudletFinish()method.Step 2: Add Retry Logic to the Broker
Modify the broker to retry failed cloudlets up to a predefined limit. You’ll need to extend theCloudletclass (or create a custom subclass) to track retry attempts. Here’s a simplified code snippet:
// First, extend Cloudlet to add retry tracking public class RetryEnabledCloudlet extends Cloudlet { private int retryCount = 0; public static final int MAX_RETRIES = 3; // Constructor and other methods... public int getRetryCount() { return retryCount; } public void incrementRetryCount() { retryCount++; } } // Then modify your custom DatacenterBroker @Override protected void onCloudletFinish(List<Cloudlet> finishedCloudlets) { for (Cloudlet cloudlet : finishedCloudlets) { if (cloudlet.getStatus() == Cloudlet.FAILED) { RetryEnabledCloudlet retryCloudlet = (RetryEnabledCloudlet) cloudlet; if (retryCloudlet.getRetryCount() < RetryEnabledCloudlet.MAX_RETRIES) { retryCloudlet.incrementRetryCount(); // Resubmit the cloudlet (you can choose the same or a new VM) submitCloudlet(retryCloudlet); System.out.printf("Retrying cloudlet %d (Attempt %d/%d)%n", cloudlet.getId(), retryCloudlet.getRetryCount(), RetryEnabledCloudlet.MAX_RETRIES); } else { System.out.printf("Cloudlet %d failed after %d retries%n", cloudlet.getId(), RetryEnabledCloudlet.MAX_RETRIES); } } } }
- Step 3: Simulate Failures for Testing
To validate your setup, you can inject failures by modifying theDatacenterorVmclasses to randomly mark cloudlets as failed, or use CloudSim’s built-in fault simulation tools if available. No optimization logic is needed here— this is a fixed-retry strategy based purely on failure detection.
To give you precise guidance, I’d love to see your current algorithm’s code or a detailed breakdown of its logic (e.g., how it decides when to retry, what failure types it handles). But here’s a framework to assess its validity and make targeted adjustments:
Is the Algorithm Reasonable?
- Does it match your failure model? If your environment mostly has transient failures (e.g., temporary VM overload), a fixed retry count works. For permanent failures (e.g., VM crashes), you should skip retrying on the same VM and reschedule to a new one immediately.
- Does it avoid "retry storms"? If multiple cloudlets fail at once, naive retries can flood the datacenter. Adding a small random delay between retries will help reduce congestion.
Does It Fit Your Requirements?
- If low latency is critical: Prioritize retries on underutilized VMs instead of reusing the failed one.
- If resource efficiency matters: Set a retry budget (e.g., total CPU time spent on failed attempts) and stop retrying if it’s exceeded.
Adjustments to Consider
- Failure-Type Specific Retries: Check why a cloudlet failed (e.g.,
FAILED_RESOURCE_UNAVAILABLEvs.FAILED_BY_VM) and tailor the strategy— retry on a new VM for crashes, but wait a short time for resource overloads. - Exponential Backoff: Instead of fixed delays, increase wait time between retries (1s → 2s → 4s) to reduce datacenter load.
- Dynamic Retry Limits: Adjust max retries based on current datacenter load— fewer retries during peak hours to avoid worsening congestion.
- Failure-Type Specific Retries: Check why a cloudlet failed (e.g.,
Validation Tips
- Run simulations with varying failure rates (10%, 30%, 50%) and compare metrics like cloudlet success rate, average completion time, and resource utilization.
- Test against a baseline (no retries) to measure how much your algorithm improves fault tolerance.
内容的提问来源于stack exchange,提问作者Fahad Khan

