You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Gatling/Artillery对AWS StepFunction进行异步负载测试?

Great question—testing asynchronous systems like AWS Step Functions does add a layer of complexity compared to synchronous APIs, but there are proven patterns you can implement with Gatling or Artillery to get meaningful results. Let’s break this down:

Core Testing Approach

Since Step Functions are async, your load test needs to cover two critical dimensions:

  1. Trigger throughput: How well your system handles concurrent StartExecution calls (QPS, success rate, throttling events)
  2. Execution reliability: The success rate, total duration, and error patterns of the actual Step Function workflows

The workflow looks like this:

  • Batch-trigger executions with your load tool
  • Track each execution’s status until it completes
  • Aggregate results to measure both trigger and execution performance

1. Implementation with Gatling

Gatling’s Scala DSL lets you chain requests and add custom logic for polling execution status. Here’s a practical example:

First, make sure you handle AWS Signature Version 4 (use the gatling-aws plugin to avoid manual signature work). Then define your scenario:

import io.gatling.core.Predef._
import io.gatling.http.Predef._
import scala.concurrent.duration._

class StepFunctionsLoadTest extends Simulation {
  val httpProtocol = http
    .baseUrl("https://states.us-east-1.amazonaws.com")
    .header("Content-Type", "application/json")
    .sign(AwsV4Signature("us-east-1", "states")) // Handled via gatling-aws plugin

  val targetStateMachineArn = "arn:aws:states:us-east-1:123456789012:stateMachine:MyProductionStateMachine"

  val scn = scenario("Step Functions End-to-End Load Test")
    // Generate a unique execution name to avoid conflicts
    .exec(session => session.set("executionName", s"load-test-${java.util.UUID.randomUUID()}"))
    // Trigger the Step Function execution
    .exec(
      http("Start Execution")
        .post("/")
        .body(StringBody(s"""{"stateMachineArn": "$targetStateMachineArn", "name": "${executionName}"}""")).asJson
        .check(jsonPath("$.executionArn").saveAs("executionArn"))
    )
    // Initial pause to avoid immediate polling
    .pause(1 second)
    // Poll until execution completes or max attempts are hit
    .repeat(12, "pollCount") {
      exec(
        http("Check Execution Status")
          .post("/")
          .body(StringBody(s"""{"executionArn": "${executionArn}"}""")).asJson
          .check(jsonPath("$.status").saveAs("executionStatus"))
      )
      // Wait only if execution is still running
      .doIf(session => !List("SUCCEEDED", "FAILED").contains(session("executionStatus").as[String])) {
        pause(5 seconds)
      }
    }
    // Log failures for debugging
    .exec(session => {
      val status = session("executionStatus").as[String]
      if (status != "SUCCEEDED") {
        println(s"Execution failed: ${session("executionArn").as[String]} | Status: $status")
      }
      session
    })

  // Configure load profile (adjust based on your test goals)
  setUp(
    scn.inject(
      rampUsers(200) during (30 seconds),
      constantUsersPerSec(20) during (2 minutes)
    ).protocols(httpProtocol)
  )
}

Key notes:

  • Adjust the repeat count and pause duration to match your Step Function’s average execution time
  • Use Gatling’s built-in metrics to track StartExecution latency and success rate, plus custom logs for execution outcomes

2. Implementation with Artillery

Artillery’s JavaScript processor support makes it easy to add async polling logic. Here’s a setup example:

First, your artillery.yml config:

config:
  target: "https://states.us-east-1.amazonaws.com"
  phases:
    - duration: 120
      arrivalRate: 15
  headers:
    Content-Type: "application/json"
  processors:
    - "./sf-test-processors.js"
scenarios:
  - name: "Step Function Load Test"
    flow:
      - function: "generateUniqueName"
      - post:
          url: "/"
          json:
            stateMachineArn: "arn:aws:states:us-east-1:123456789012:stateMachine:MyProductionStateMachine"
            name: "{{ executionName }}"
          capture:
            json: "$.executionArn"
            as: "executionArn"
      - function: "pollForCompletion"

Then the corresponding sf-test-processors.js file (uses the AWS SDK for simplified API calls):

const AWS = require('aws-sdk');
const stepFunctions = new AWS.StepFunctions({ region: 'us-east-1' });

module.exports = {
  generateUniqueName: (context, events, done) => {
    context.vars.executionName = `load-test-${Math.random().toString(36).slice(2, 15)}`;
    done();
  },
  pollForCompletion: async (context, events, done) => {
    const params = { executionArn: context.vars.executionArn };
    const maxAttempts = 10;
    const pollInterval = 5000; // 5 seconds

    for (let attempt = 0; attempt < maxAttempts; attempt++) {
      const result = await stepFunctions.describeExecution(params).promise();
      if (result.status === 'SUCCEEDED' || result.status === 'FAILED') {
        if (result.status !== 'SUCCEEDED') {
          context.fail(`Execution failed: ${context.vars.executionArn} | Status: ${result.status}`);
        }
        return done();
      }
      await new Promise(resolve => setTimeout(resolve, pollInterval));
    }

    context.fail(`Execution timed out: ${context.vars.executionArn}`);
    done();
  }
};

Key notes:

  • Artillery automatically tracks request metrics, and custom failures will show up in your test report
  • Use AWS IAM roles (or environment variables) to authenticate the SDK client

3. Optimizations for High Concurrency

If you’re testing at scale, naive polling can put extra load on the Step Functions API. Instead:

  • Use CloudWatch Events + SQS: Configure your Step Function to send state change events to an SQS queue. Your load tool only triggers executions, then consumes results from SQS asynchronously. This reduces polling overhead drastically.
  • Separate metrics: Track StartExecution performance (throttling, latency) separately from Step Function execution metrics (success rate, total duration) to isolate bottlenecks.
  • Handle throttling: Build retry logic for StartExecution calls (AWS returns ThrottlingException when you hit rate limits) to get accurate throughput numbers.

内容的提问来源于stack exchange,提问作者Theo Cerutti

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 15:02:53