F#中带限流的批量并发HTTP GET调用队列模式问询
Absolutely! F# has all the tools you need to build a throttled, queue-based concurrent HTTP caller that fits your requirements perfectly. Let's walk through how to implement this—including handling both successful responses and errors like 404s or 500s.
First, Define a Result Type
We'll start with a union type to capture both successful results and failures, so we can collect all outcomes without losing data about which URL succeeded or failed:
type HttpCallResult<'T> = | Success of url: string * result: 'T | Failure of url: string * error: exn
Throttled Concurrent Implementation
The core of this uses asynchronous workflows (F#'s Async) and a SemaphoreSlim to enforce the concurrency limit. This setup runs requests continuously as slots free up (instead of strict batches, though we'll cover that alternative too):
open System open System.Net.Http open System.Threading /// Executes HTTP calls with a max concurrent limit, returning all success/failure results let throttledHttpCalls (maxConcurrent: int) (callFunc: string -> Async<'T>) (urls: string list) : Async<HttpCallResult<'T> list> = async { // Semaphore controls how many requests can run at the same time use semaphore = new SemaphoreSlim(maxConcurrent, maxConcurrent) let processSingleUrl url = async { // Wait for an available slot in the semaphore do! semaphore.WaitAsync() |> Async.AwaitTask try // Execute the custom HTTP call function provided let! result = callFunc url return Success(url, result) with ex -> // Capture any errors (HTTP failures, timeouts, network issues, etc.) return Failure(url, ex) finally // Always release the semaphore slot, even if the request failed semaphore.Release() |> ignore } // Create async tasks for all URLs, then run them in parallel (semaphore enforces limits) let allTasks = urls |> List.map processSingleUrl return! Async.Parallel allTasks |> Async.map Array.toList }
Key Details to Note
- SemaphoreSlim: This is the secret to throttling. It ensures only
maxConcurrentrequests are active at any time, preventing you from overwhelming the target server or hitting rate limits. - Async Workflows: F#'s Async is ideal here because it's lightweight and optimized for I/O-bound tasks like HTTP calls.
Async.Parallelmanages running the tasks, while the semaphore keeps concurrency in check. - Flexible Call Function: The
callFuncparameter lets you pass in any custom HTTP logic—whether you're fetching HTML, parsing JSON, adding auth headers, or handling timeouts. You're not locked into a specific implementation. - Error Handling: Any exceptions (including non-2xx HTTP status codes if you use
EnsureSuccessStatusCode()) get captured asFailurecases, so you can inspect them later without losing track of which URL failed.
Example: Fetching HTML Content
Here's how you'd use this to fetch HTML from a list of URLs with a max of 5 concurrent requests:
let fetchHtmlFromUrls urls = async { // Reuse HttpClient (critical to avoid socket exhaustion—never create one per request!) use httpClient = new HttpClient() let fetchSingleUrl url = async { let! response = httpClient.GetAsync(url) |> Async.AwaitTask // Throw an exception for non-2xx status codes (caught in the Failure case) response.EnsureSuccessStatusCode() |> ignore let! html = response.Content.ReadAsStringAsync() |> Async.AwaitTask return html } // Run with max 5 concurrent requests let! results = throttledHttpCalls 5 fetchSingleUrl urls return results } // To execute: // let myUrls = ["https://example.com"; "https://fsharp.org"; "https://github.com"] // let results = fetchHtmlFromUrls myUrls |> Async.RunSynchronously
Alternative: Strict Batch Processing
If you need strict batch behavior (wait for all 5 requests to finish before starting the next 5), you can split the URLs into chunks and process each batch sequentially:
let batchHttpCalls (batchSize: int) (callFunc: string -> Async<'T>) (urls: string list) : Async<HttpCallResult<'T> list> = async { let batches = urls |> List.chunkBySize batchSize let rec processBatches remainingBatches accumulatedResults = async { match remainingBatches with | [] -> return List.rev accumulatedResults | currentBatch :: rest -> // Process all requests in the current batch in parallel let! batchResults = currentBatch |> List.map (fun url -> async { try let! result = callFunc url return Success(url, result) with ex -> return Failure(url, ex) }) |> Async.Parallel |> Async.map Array.toList // Add batch results to the accumulator and process the next batch return! processBatches rest (batchResults @ accumulatedResults) } return! processBatches batches [] }
Final Tips
- Reuse HttpClient: Never create a new
HttpClientper request—this leads to socket exhaustion. Always reuse a single instance (as shown in the example). - Customize Error Handling: If you don't want to treat non-2xx codes as failures, remove
EnsureSuccessStatusCode()and instead checkresponse.StatusCodeexplicitly, then include it in your result type for granular handling.
内容的提问来源于stack exchange,提问作者App2015

