如何并行化Python API调用?Spotify艺人新作品监控程序提速需求
Great question—parallelizing I/O-bound tasks like API calls is exactly the way to slash that runtime for your thousands of artists. Let’s walk through the most practical approaches in Python, tailored to your Spotify workflow.
Method 1: ThreadPoolExecutor (Simple & Effective for I/O-Bound Tasks)
Since API calls spend most of their time waiting for network responses (not CPU work), using threads is perfect here—Python’s GIL doesn’t block I/O-bound operations. This approach is easy to implement with the built-in concurrent.futures module.
First, wrap your per-artist logic into a reusable function, then run it across a thread pool:
import requests from concurrent.futures import ThreadPoolExecutor, as_completed # Replace with your actual Spotify auth headers YOUR_AUTH_HEADERS = {"Authorization": "Bearer YOUR_ACCESS_TOKEN"} def process_single_artist(artist_id): try: # Step 1: Validate the artist exists on Spotify artist_resp = requests.get( f"https://api.spotify.com/v1/artists/{artist_id}", headers=YOUR_AUTH_HEADERS ) artist_resp.raise_for_status() # Trigger error for 4xx/5xx responses # Step 2: Fetch album count albums_resp = requests.get( f"https://api.spotify.com/v1/artists/{artist_id}/albums", headers=YOUR_AUTH_HEADERS, params={"limit": 50} # Adjust limit if you need more than 50 albums ) albums_resp.raise_for_status() album_count = len(albums_resp.json()["items"]) return (artist_id, album_count, None) except Exception as e: # Handle errors like missing artists/albums gracefully return (artist_id, None, str(e)) # Your list of 1000+ artist IDs artist_ids = ["artist_id_1", "artist_id_2", ...] # Run tasks in parallel results = [] # Start with max_workers=20 (adjust based on Spotify's rate limits) with ThreadPoolExecutor(max_workers=20) as executor: # Submit all tasks to the pool futures = [executor.submit(process_single_artist, aid) for aid in artist_ids] # Process results as they complete (instead of waiting for all) for future in as_completed(futures): results.append(future.result()) # Now compare results with your CSV data for artist_id, count, error in results: if error: print(f"Failed to process {artist_id}: {error}") else: # Add your CSV comparison/update logic here pass
Method 2: Async I/O with aiohttp (More Efficient for High Volume)
For even better performance with thousands of requests, use async non-blocking I/O. This uses a single thread but handles multiple requests concurrently, making it ideal for large batches.
import aiohttp import asyncio YOUR_AUTH_HEADERS = {"Authorization": "Bearer YOUR_ACCESS_TOKEN"} async def process_artist_async(session, artist_id): try: # Validate artist async with session.get( f"https://api.spotify.com/v1/artists/{artist_id}", headers=YOUR_AUTH_HEADERS ) as resp: resp.raise_for_status() await resp.json() # Just confirm the artist exists # Fetch album count async with session.get( f"https://api.spotify.com/v1/artists/{artist_id}/albums", headers=YOUR_AUTH_HEADERS, params={"limit": 50} ) as resp: resp.raise_for_status() albums_data = await resp.json() album_count = len(albums_data["items"]) return (artist_id, album_count, None) except Exception as e: return (artist_id, None, str(e)) async def main(): artist_ids = ["artist_id_1", "artist_id_2", ...] # Use a semaphore to avoid hitting Spotify's rate limits semaphore = asyncio.Semaphore(20) # Wrap the process function to respect the semaphore async def bounded_process(aid): async with semaphore: return await process_artist_async(session, aid) # Create a single HTTP session (more efficient than multiple sessions) async with aiohttp.ClientSession() as session: # Launch all tasks and wait for results tasks = [bounded_process(aid) for aid in artist_ids] results = await asyncio.gather(*tasks) # Process results for artist_id, count, error in results: if error: print(f"Failed {artist_id}: {error}") else: # Add CSV comparison logic here pass if __name__ == "__main__": asyncio.run(main())
Critical Things to Keep in Mind
- Respect Spotify’s Rate Limits: Spotify’s free tier allows ~180 requests/minute, premium tiers up to 1500. Adjust
max_workers(ThreadPool) orSemaphore(async) to stay under this. Check theX-RateLimit-RemainingandX-RateLimit-Resetheaders in responses to dynamically adjust if needed. - Error Handling: The try/except blocks catch issues like missing artists or albums—you can refine this to handle specific errors (e.g., 404s for missing artists) differently.
- Authentication: Never hardcode your access token—use environment variables or a secure config file.
- Caching: If you run this daily, cache valid artist IDs (e.g., in a JSON file or small database) to avoid re-validating them every time.
内容的提问来源于stack exchange,提问作者1.23423E+14

