如何实现自动遍历带分页参数的URL并快速检测指定字符串?
Building a High-Speed Page Checker for Paged URLs
Core Approach
To hit the minimum 5 requests per second requirement, sequential requests won’t cut it—we need asynchronous concurrency. Python’s aiohttp library is perfect here: it lets us fire off multiple requests at once without blocking, which easily meets the speed target while keeping the code clean.
The plan breaks down into 4 simple steps:
- Generate all paged URLs from page 1 to
maxPageNumber - Asynchronously fetch each page’s content
- Check if your target string (e.g., "and") exists in the response
- Track results (which pages have the string, and any failed requests)
Working Code Example
Here’s a complete, runnable script that implements this logic:
import asyncio import aiohttp from typing import List, Tuple async def check_page(session: aiohttp.ClientSession, url: str, target_str: str) -> Tuple[str, bool, str]: """Fetch a page and check if the target string exists.""" try: async with session.get(url, timeout=10) as response: response.raise_for_status() # Trigger error for HTTP 4xx/5xx codes content = await response.text() has_target = target_str in content return (url, has_target, "Success") except Exception as e: return (url, False, f"Failed: {str(e)}") async def main(base_url: str, max_page: int, target_str: str, concurrency_limit: int = 10): # Generate all paged URLs (adjust the placeholder to match your site's pattern) urls = [base_url.format(page) for page in range(1, max_page + 1)] # Use a semaphore to control concurrency (prevents overwhelming the server) semaphore = asyncio.Semaphore(concurrency_limit) async with aiohttp.ClientSession() as session: # Wrap the check function to respect the concurrency limit async def bounded_check(url): async with semaphore: return await check_page(session, url, target_str) # Run all checks at once tasks = [bounded_check(url) for url in urls] results = await asyncio.gather(*tasks) # Print a clean summary print(f"Finished checking {len(results)} pages:\n") for url, has_target, status in results: if has_target: print(f"✅ Found '{target_str}' in {url}") elif "Failed" in status: print(f"❌ {status} for {url}") else: print(f"🔍 '{target_str}' not found in {url}") if __name__ == "__main__": # Configure these values for your use case BASE_URL = "https://example.com/articles/page/{}" # Replace with your site's pagination URL MAX_PAGE_NUMBER = 50 TARGET_STRING = "and" CONCURRENCY_LIMIT = 10 # Higher = faster, but don't spam the server # Launch the async program asyncio.run(main(BASE_URL, MAX_PAGE_NUMBER, TARGET_STRING, CONCURRENCY_LIMIT))
Key Tips for Speed & Reliability
- Concurrency Tuning: A
concurrency_limitof 10 will easily exceed 5 requests per second. If the server tolerates it, bump this number up to go even faster. - Session Reuse: Using a single
ClientSessionreuses TCP connections, which cuts down on overhead and speeds up requests drastically. - Error Resilience: The script catches timeouts, HTTP errors, and network blips, so a single failed page won’t crash the whole process.
- Ethical Scraping: Always check the site’s
robots.txtand terms of service. If you hit rate limits, add small delays (but with concurrency, you shouldn’t need this to meet the 5/sec target).
Verify Your Speed
To confirm you’re hitting the rate requirement, add timing code to the main() function:
start_time = asyncio.get_event_loop().time() # ... run tasks ... elapsed = asyncio.get_event_loop().time() - start_time print(f"\nTotal time: {elapsed:.2f} seconds") print(f"Requests per second: {len(results)/elapsed:.2f}")
This will show exactly how fast your checker is running.
内容的提问来源于stack exchange,提问作者Simplex1
相关产品推荐
相关产品推荐

