PuppeteerSharp取消大文件下载遇问题:IDM弹窗与Content-Length缺失
Let’s break down your two problems and fix them step by step:
1. Fixing IDM Popup (--disable-extensions Not Working)
The --disable-extensions flag doesn’t block IDM because IDM isn’t a regular Chrome extension—it hooks into the browser at a system level (via BHOs or plugin integrations). To fully prevent IDM from intercepting Puppeteer’s downloads, you need to add more aggressive flags and isolate the browser profile:
Updated Launch Configuration
var downloadPath = Path.Combine(Environment.GetFolderPath(Environment.SpecialFolder.ApplicationData), "Puppeteer_Downloads"); Directory.CreateDirectory(downloadPath); // Ensure the download folder exists first var browser = await Puppeteer.LaunchAsync(new LaunchOptions { Headless = true, ExecutablePath = ChromePath, IgnoreHTTPSErrors = true, Args = new[] { "--disable-extensions", "--disable-plugins", "--disable-plugin-extensions", "--disable-background-downloads", "--disable-features=DownloadBubble,DownloadBubbleV2", // Disable Chrome's native download UI "--no-default-browser-check", "--user-data-dir=./temp_puppeteer_profile" // Use a fresh, isolated profile to avoid IDM hooks } }); // Ensure Puppeteer fully controls downloads await page.Client.SendAsync("Page.setDownloadBehavior", new { behavior = "allow", downloadPath = downloadPath, allowOverwrite = true // Optional: Overwrite existing files });
Why This Works:
- The extra flags disable plugins and background download handlers that IDM relies on.
- Using a temporary user directory ensures Puppeteer doesn’t load any existing browser configurations that might have IDM integrations.
- Explicitly setting
downloadPathand enabling overwrites ensures Puppeteer takes full ownership of downloads.
2. Detecting File Size When Content-Length Is Missing in Requests
You’re checking the request headers for Content-Length, but that’s the size of the data sent to the server—not the file you’re downloading. The file size lives in the response headers sent by the server. Plus, some servers use chunked transfer encoding (no Content-Length), so we need a fallback approach.
Solution: Use HEAD Requests to Pre-Check File Size
Before allowing the download request to proceed, send a HEAD request to the server to fetch the file size without downloading the entire file. If it exceeds your limit, abort the original request.
await page.SetRequestInterceptionAsync(true); page.Request += async (sender, e) => { var request = e.Request; // Target only downloadable resources (adjust filters to match your use case) var isDownloadable = request.ResourceType == ResourceType.Other || request.Url.EndsWith(".zip") || request.Url.EndsWith(".exe") || request.Url.EndsWith(".pdf"); if (isDownloadable) { try { // Send a HEAD request to get file size without full download using var httpClient = new HttpClient(); var headRequest = new HttpRequestMessage(HttpMethod.Head, request.Url); // Copy original request headers (e.g., auth cookies) to access restricted resources foreach (var header in request.Headers) { if (!headRequest.Headers.TryAddWithoutValidation(header.Key, header.Value)) { headRequest.Content?.Headers.TryAddWithoutValidation(header.Key, header.Value); } } var headResponse = await httpClient.SendAsync(headRequest); if (headResponse.IsSuccessStatusCode) { // Check if Content-Length is available if (headResponse.Content.Headers.ContentLength.HasValue) { var fileSize = headResponse.Content.Headers.ContentLength.Value; if (fileSize > GIVENSIZE) { await request.AbortAsync(); Console.WriteLine($"Aborted download: {request.Url} (size exceeds limit)"); return; } } else { // Handle chunked transfer (no Content-Length) // Option 1: Allow request and monitor download progress (more complex) // Option 2: Abort if size can't be verified // await request.AbortAsync(); // Console.WriteLine($"Aborted download: {request.Url} (no Content-Length available)"); // return; } } } catch (Exception ex) { // Handle cases where HEAD requests aren't supported by the server Console.WriteLine($"Failed to check file size: {ex.Message}"); } } // Allow request to proceed if all checks pass await request.ContinueAsync(); };
Key Notes:
- Copying headers from the original request ensures you can access authenticated or restricted resources.
- For servers that don’t support
HEADrequests, adjust the logic to either allow the download (and monitor progress if needed) or abort it based on your requirements. - If you need to handle chunked transfers, you’d need to listen to the
Responseevent and track downloaded bytes, then terminate the request if it exceeds your limit—this is more involved but covers edge cases.
内容的提问来源于stack exchange,提问作者Inside Man

