Haskell学习练习:迁移Elixir批量网页爬取项目,实现带1秒延迟的2500个Zip码页面爬取
Hey there! Since you're already comfortable with single-page scraping using Scalpel, extending this to batch processing with a 1-second delay between requests is straightforward once you wrap your head around how Haskell handles sequential IO actions. Let's walk through this step by step.
Core Approach
The key here is to create a reusable IO action for scraping a single zip code (including the delay), then execute this action sequentially for every zip code in your list. Haskell's mapM_ function is perfect for this—it takes a list of values, applies an IO function to each, and runs all the resulting IO actions in order.
Step 1: Set Up Dependencies
First, make sure your Cabal/Stack file includes the necessary packages:
scalpelfor scrapingbase(forControl.ConcurrentandControl.Exception)http-clientandhttp-client-tls(Scalpel relies on these for HTTP requests)
Step 2: Write the Single-Zip Scraping Function
This function will handle constructing the URL, scraping the page, processing results, and adding the 1-second delay. We'll also add basic error handling to avoid crashing the entire batch if one request fails.
import Text.HTML.Scalpel import Control.Concurrent (threadSleep) import Control.Exception (catch, HttpException(..)) import Network.HTTP.Client (HttpExceptionContent(StatusCodeException), responseStatus) -- Define your scraper selector here (adjust to match the data you want) zipInfoSelector :: Scraper String String zipInfoSelector = text $ "div" @: [hasClass "zip-details"] scrapeZip :: String -> IO () scrapeZip zipCode = do let url = "http://www.acme.org/zip-info?zip=" ++ zipCode putStrLn $ "Scraping zip code: " ++ zipCode -- Attempt to scrape the page, with error handling result <- scrapeURL url zipInfoSelector `catch` handleScrapeError zipCode -- Process the result (print, save to file, etc.) case result of Just info -> putStrLn $ "Successfully scraped: " ++ info Nothing -> putStrLn $ "No data found for zip " ++ zipCode -- Wait 1 second (note: threadSleep takes microseconds) threadSleep 1000000 where -- Handle HTTP errors gracefully handleScrapeError :: String -> HttpException -> IO (Maybe String) handleScrapeError z e = do case e of HttpExceptionRequest _ (StatusCodeException resp _) -> putStrLn $ "Error scraping " ++ z ++ ": HTTP status " ++ show (responseStatus resp) _ -> putStrLn $ "Error scraping " ++ z ++ ": " ++ show e return Nothing
Step 3: Execute the Batch in Main
In your main function, you just need to pass your list of zip codes to mapM_ scrapeZip. You can hardcode the list, read it from a file, or generate it programmatically.
main :: IO () main = do -- Example zip code list (replace with your full 2500 zips) let zipCodes = ["90210", "10001", "60601"] -- Extend this to your complete list -- Run the scraper for each zip code sequentially mapM_ scrapeZip zipCodes putStrLn "Batch scraping complete!"
Key Notes
- Sequential Execution:
mapM_runs eachscrapeZipaction one after another, which ensures the 1-second delay is respected between every request—exactly what you want to avoid overwhelming the target server. - Thread Sleep Units: Remember that
threadSleepuses microseconds, so 1 second is1000000(not1). - Error Handling: The
catchblock ensures that if one request fails (e.g., 404, network issue), the program will log the error and continue with the next zip code instead of crashing. - No Concurrency: Since you don't need speed, sequential execution is ideal. If you ever wanted to add controlled concurrency later, you could use libraries like
async, but that's unnecessary here.
Why This Works
In Haskell, IO actions are first-class values—scrapeZip zipCode returns an IO action that encapsulates all the side effects (HTTP request, printing, sleeping). mapM_ takes a list of these actions and chains them together into a single IO action that runs in order. This aligns perfectly with your original Elixir logic of iterating through the list with a delay between each step.
内容的提问来源于stack exchange,提问作者Jeroen Bourgois

