You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现自动遍历带分页参数的URL并快速检测指定字符串?

Building a High-Speed Page Checker for Paged URLs

Core Approach

To hit the minimum 5 requests per second requirement, sequential requests won’t cut it—we need asynchronous concurrency. Python’s aiohttp library is perfect here: it lets us fire off multiple requests at once without blocking, which easily meets the speed target while keeping the code clean.

The plan breaks down into 4 simple steps:

  1. Generate all paged URLs from page 1 to maxPageNumber
  2. Asynchronously fetch each page’s content
  3. Check if your target string (e.g., "and") exists in the response
  4. Track results (which pages have the string, and any failed requests)

Working Code Example

Here’s a complete, runnable script that implements this logic:

import asyncio
import aiohttp
from typing import List, Tuple

async def check_page(session: aiohttp.ClientSession, url: str, target_str: str) -> Tuple[str, bool, str]:
    """Fetch a page and check if the target string exists."""
    try:
        async with session.get(url, timeout=10) as response:
            response.raise_for_status()  # Trigger error for HTTP 4xx/5xx codes
            content = await response.text()
            has_target = target_str in content
            return (url, has_target, "Success")
    except Exception as e:
        return (url, False, f"Failed: {str(e)}")

async def main(base_url: str, max_page: int, target_str: str, concurrency_limit: int = 10):
    # Generate all paged URLs (adjust the placeholder to match your site's pattern)
    urls = [base_url.format(page) for page in range(1, max_page + 1)]
    
    # Use a semaphore to control concurrency (prevents overwhelming the server)
    semaphore = asyncio.Semaphore(concurrency_limit)
    
    async with aiohttp.ClientSession() as session:
        # Wrap the check function to respect the concurrency limit
        async def bounded_check(url):
            async with semaphore:
                return await check_page(session, url, target_str)
        
        # Run all checks at once
        tasks = [bounded_check(url) for url in urls]
        results = await asyncio.gather(*tasks)
        
        # Print a clean summary
        print(f"Finished checking {len(results)} pages:\n")
        for url, has_target, status in results:
            if has_target:
                print(f"✅ Found '{target_str}' in {url}")
            elif "Failed" in status:
                print(f"❌ {status} for {url}")
            else:
                print(f"🔍 '{target_str}' not found in {url}")

if __name__ == "__main__":
    # Configure these values for your use case
    BASE_URL = "https://example.com/articles/page/{}"  # Replace with your site's pagination URL
    MAX_PAGE_NUMBER = 50
    TARGET_STRING = "and"
    CONCURRENCY_LIMIT = 10  # Higher = faster, but don't spam the server
    
    # Launch the async program
    asyncio.run(main(BASE_URL, MAX_PAGE_NUMBER, TARGET_STRING, CONCURRENCY_LIMIT))

Key Tips for Speed & Reliability

  • Concurrency Tuning: A concurrency_limit of 10 will easily exceed 5 requests per second. If the server tolerates it, bump this number up to go even faster.
  • Session Reuse: Using a single ClientSession reuses TCP connections, which cuts down on overhead and speeds up requests drastically.
  • Error Resilience: The script catches timeouts, HTTP errors, and network blips, so a single failed page won’t crash the whole process.
  • Ethical Scraping: Always check the site’s robots.txt and terms of service. If you hit rate limits, add small delays (but with concurrency, you shouldn’t need this to meet the 5/sec target).

Verify Your Speed

To confirm you’re hitting the rate requirement, add timing code to the main() function:

start_time = asyncio.get_event_loop().time()
# ... run tasks ...
elapsed = asyncio.get_event_loop().time() - start_time
print(f"\nTotal time: {elapsed:.2f} seconds")
print(f"Requests per second: {len(results)/elapsed:.2f}")

This will show exactly how fast your checker is running.

内容的提问来源于stack exchange,提问作者Simplex1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:42:38