如何使用requests.Session对指定URL实现请求速率限制,确保连续请求间的最小延迟
Absolutely, you can implement request throttling directly on a requests.Session object—either by subclassing the session (the cleanest approach) or using event hooks combined with wrapping the session's request logic. Both methods meet your requirement of modifying the session itself, so let's break them down with code examples.
Subclassing requests.Session (Recommended)
Subclassing Session lets you encapsulate all throttling logic in a single class, making it easy to reuse and maintain. We'll override the send method (which all request types like get, post, etc., ultimately call) to add delay checks before each request:
import time from requests import Session class ThrottledSession(Session): def __init__(self, min_delay: float): super().__init__() self.min_delay = min_delay # Minimum seconds between requests self.last_request_time = 0 def send(self, request, **kwargs): # Calculate time elapsed since last request time_since_last = time.time() - self.last_request_time # Sleep if we haven't waited the full minimum delay if time_since_last < self.min_delay: time.sleep(self.min_delay - time_since_last) # Update the timestamp before sending the request self.last_request_time = time.time() # Pass through to the original send method return super().send(request, **kwargs)
How to use it:
# Create a session with a 2-second minimum delay between requests session = ThrottledSession(min_delay=2) # All requests through this session will respect the throttle session.get("https://example.com") # Next request will automatically wait until 2 seconds have passed session.post("https://example.com/submit", data={"key": "value"})
Using Event Hooks (Alternative Approach)
If you specifically want to leverage requests' built-in event hooks, you can combine a response hook to track request completion times with a wrapped request method to enforce delays. Here's how:
import time from requests import Session def throttle_response_hook(response, *args, **kwargs): # Record when the response was received (marks the end of the request) response.session.last_request_time = time.time() return response def create_throttled_session(min_delay: float) -> Session: session = Session() session.min_delay = min_delay session.last_request_time = 0 # Attach the hook to track request completion session.hooks["response"].append(throttle_response_hook) # Wrap the session's original request method to add throttling original_request = session.request def throttled_request(method, url, **kwargs): elapsed = time.time() - session.last_request_time if elapsed < session.min_delay: time.sleep(session.min_delay - elapsed) return original_request(method, url, **kwargs) session.request = throttled_request return session
How to use it:
session = create_throttled_session(min_delay=1.5) session.get("https://example.com") # The next request will wait until 1.5 seconds have passed since the first completed session.get("https://example.com/another")
Thread-Safe Throttling
If you're using the session in a multi-threaded environment, the basic implementations above can have race conditions (multiple threads updating last_request_time at the same time). To fix this, add a lock:
import time import threading from requests import Session class ThreadSafeThrottledSession(Session): def __init__(self, min_delay: float): super().__init__() self.min_delay = min_delay self.last_request_time = 0 self._lock = threading.Lock() def send(self, request, **kwargs): with self._lock: time_since_last = time.time() - self.last_request_time if time_since_last < self.min_delay: time.sleep(self.min_delay - time_since_last) self.last_request_time = time.time() return super().send(request, **kwargs)
Key Notes
- Subclassing is preferred because it keeps all logic contained and is easier to extend (e.g., adding retry logic later).
- The hook-based method works but requires wrapping the
requestmethod, which is a bit more indirect. - Always consider thread safety if you're using the session across multiple threads.
内容的提问来源于stack exchange,提问作者Peteris

