如何用aiohttp异步批量下载大量PDF文件?(Python异步新手求助)
Hey there! I totally get where you're coming from—async code can feel like a whole new world when you're starting out, especially with aiohttp. Let's take your working single-URL code and expand it into a clean, understandable batch download solution that runs in a single thread.
Core Idea
We'll wrap your single download logic into a reusable async function, then use asyncio.gather() to run all these download tasks concurrently. This keeps everything in one thread and leverages async I/O to handle multiple downloads without waiting for each one to finish before starting the next.
Step-by-Step Implementation
1. Wrap the Single Download Logic into a Function
First, let's turn your existing code into a standalone async function. We'll add error handling so a failed download doesn't break the whole batch, and let it accept a URL and output filename for flexibility.
import aiohttp import aiofiles import asyncio async def download_pdf(session: aiohttp.ClientSession, url: str, save_path: str): try: async with session.get(url) as resp: if resp.status == 200: # Open the file asynchronously to avoid blocking the event loop async with aiofiles.open(save_path, mode='wb') as f: # Write the PDF content in one go (great for small-to-medium files) await f.write(await resp.read()) print(f"Successfully saved: {save_path}") else: print(f"Failed to fetch {url}: Status code {resp.status}") except Exception as e: print(f"Error processing {url}: {str(e)}")
2. Prepare Your List of URLs
Create a list of your 100 PDF URLs. For example, here's a sample structure (replace this with your actual URLs):
pdf_urls = [ "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf", # Add 99 more URLs here... "https://example.com/another-sample.pdf" ]
To avoid overwriting files, map each URL to a unique filename (like pdf_001.pdf, pdf_002.pdf):
# Create a list of (url, filename) tuples for batch processing download_jobs = [ (url, f"downloaded_pdf_{index+1:03d}.pdf") for index, url in enumerate(pdf_urls) ]
3. Run All Downloads Concurrently
Now we'll create the main async function that sets up a single ClientSession (reusing one session is far more efficient than creating new ones) and runs all download tasks at once with asyncio.gather().
async def main(): # Create a single ClientSession to reuse connections across all requests async with aiohttp.ClientSession() as session: # Generate all the async download coroutines download_coroutines = [ download_pdf(session, url, save_path) for url, save_path in download_jobs ] # Execute all coroutines concurrently await asyncio.gather(*download_coroutines) if __name__ == "__main__": # Start the async event loop to run our main function asyncio.run(main())
Pro Tips for Async Newbies
- Reuse the ClientSession: Creating a new session per request wastes resources. Reusing one session keeps connections alive and speeds up batch operations.
- Error Handling Matters: The try/except block in
download_pdfensures that a single failed URL won't crash the entire batch—critical for large-scale downloads. - Stream Large PDFs: If you're dealing with huge PDFs, replace
await resp.read()with streaming chunks to save memory:async with aiofiles.open(save_path, mode='wb') as f: async for chunk in resp.content.iter_chunked(1024*1024): # 1MB chunks await f.write(chunk) - Event Loop Basics:
asyncio.run(main())starts the async event loop, which manages all concurrent tasks in a single thread—no extra threads needed!
内容的提问来源于stack exchange,提问作者Shantanu Bedajna

