You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Haskell学习练习:迁移Elixir批量网页爬取项目,实现带1秒延迟的2500个Zip码页面爬取

Batch Scraping with Rate Limiting in Haskell (Scalpel + Thread Sleep)

Hey there! Since you're already comfortable with single-page scraping using Scalpel, extending this to batch processing with a 1-second delay between requests is straightforward once you wrap your head around how Haskell handles sequential IO actions. Let's walk through this step by step.

Core Approach

The key here is to create a reusable IO action for scraping a single zip code (including the delay), then execute this action sequentially for every zip code in your list. Haskell's mapM_ function is perfect for this—it takes a list of values, applies an IO function to each, and runs all the resulting IO actions in order.

Step 1: Set Up Dependencies

First, make sure your Cabal/Stack file includes the necessary packages:

  • scalpel for scraping
  • base (for Control.Concurrent and Control.Exception)
  • http-client and http-client-tls (Scalpel relies on these for HTTP requests)

Step 2: Write the Single-Zip Scraping Function

This function will handle constructing the URL, scraping the page, processing results, and adding the 1-second delay. We'll also add basic error handling to avoid crashing the entire batch if one request fails.

import Text.HTML.Scalpel
import Control.Concurrent (threadSleep)
import Control.Exception (catch, HttpException(..))
import Network.HTTP.Client (HttpExceptionContent(StatusCodeException), responseStatus)

-- Define your scraper selector here (adjust to match the data you want)
zipInfoSelector :: Scraper String String
zipInfoSelector = text $ "div" @: [hasClass "zip-details"]

scrapeZip :: String -> IO ()
scrapeZip zipCode = do
  let url = "http://www.acme.org/zip-info?zip=" ++ zipCode
  putStrLn $ "Scraping zip code: " ++ zipCode
  
  -- Attempt to scrape the page, with error handling
  result <- scrapeURL url zipInfoSelector `catch` handleScrapeError zipCode
  
  -- Process the result (print, save to file, etc.)
  case result of
    Just info -> putStrLn $ "Successfully scraped: " ++ info
    Nothing -> putStrLn $ "No data found for zip " ++ zipCode
  
  -- Wait 1 second (note: threadSleep takes microseconds)
  threadSleep 1000000
  where
    -- Handle HTTP errors gracefully
    handleScrapeError :: String -> HttpException -> IO (Maybe String)
    handleScrapeError z e = do
      case e of
        HttpExceptionRequest _ (StatusCodeException resp _) ->
          putStrLn $ "Error scraping " ++ z ++ ": HTTP status " ++ show (responseStatus resp)
        _ -> putStrLn $ "Error scraping " ++ z ++ ": " ++ show e
      return Nothing

Step 3: Execute the Batch in Main

In your main function, you just need to pass your list of zip codes to mapM_ scrapeZip. You can hardcode the list, read it from a file, or generate it programmatically.

main :: IO ()
main = do
  -- Example zip code list (replace with your full 2500 zips)
  let zipCodes = ["90210", "10001", "60601"] -- Extend this to your complete list
  
  -- Run the scraper for each zip code sequentially
  mapM_ scrapeZip zipCodes
  
  putStrLn "Batch scraping complete!"

Key Notes

  • Sequential Execution: mapM_ runs each scrapeZip action one after another, which ensures the 1-second delay is respected between every request—exactly what you want to avoid overwhelming the target server.
  • Thread Sleep Units: Remember that threadSleep uses microseconds, so 1 second is 1000000 (not 1).
  • Error Handling: The catch block ensures that if one request fails (e.g., 404, network issue), the program will log the error and continue with the next zip code instead of crashing.
  • No Concurrency: Since you don't need speed, sequential execution is ideal. If you ever wanted to add controlled concurrency later, you could use libraries like async, but that's unnecessary here.

Why This Works

In Haskell, IO actions are first-class values—scrapeZip zipCode returns an IO action that encapsulates all the side effects (HTTP request, printing, sleeping). mapM_ takes a list of these actions and chains them together into a single IO action that runs in order. This aligns perfectly with your original Elixir logic of iterating through the list with a delay between each step.

内容的提问来源于stack exchange,提问作者Jeroen Bourgois

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.01 01:29:06