You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PuppeteerSharp取消大文件下载遇问题:IDM弹窗与Content-Length缺失

Solutions to Your Puppeteer Download Issues

Let’s break down your two problems and fix them step by step:

1. Fixing IDM Popup (--disable-extensions Not Working)

The --disable-extensions flag doesn’t block IDM because IDM isn’t a regular Chrome extension—it hooks into the browser at a system level (via BHOs or plugin integrations). To fully prevent IDM from intercepting Puppeteer’s downloads, you need to add more aggressive flags and isolate the browser profile:

Updated Launch Configuration

var downloadPath = Path.Combine(Environment.GetFolderPath(Environment.SpecialFolder.ApplicationData), "Puppeteer_Downloads");
Directory.CreateDirectory(downloadPath); // Ensure the download folder exists first

var browser = await Puppeteer.LaunchAsync(new LaunchOptions 
{ 
    Headless = true, 
    ExecutablePath = ChromePath, 
    IgnoreHTTPSErrors = true, 
    Args = new[] 
    { 
        "--disable-extensions",
        "--disable-plugins",
        "--disable-plugin-extensions",
        "--disable-background-downloads",
        "--disable-features=DownloadBubble,DownloadBubbleV2", // Disable Chrome's native download UI
        "--no-default-browser-check",
        "--user-data-dir=./temp_puppeteer_profile" // Use a fresh, isolated profile to avoid IDM hooks
    } 
});

// Ensure Puppeteer fully controls downloads
await page.Client.SendAsync("Page.setDownloadBehavior", new { 
    behavior = "allow", 
    downloadPath = downloadPath,
    allowOverwrite = true // Optional: Overwrite existing files
});

Why This Works:

  • The extra flags disable plugins and background download handlers that IDM relies on.
  • Using a temporary user directory ensures Puppeteer doesn’t load any existing browser configurations that might have IDM integrations.
  • Explicitly setting downloadPath and enabling overwrites ensures Puppeteer takes full ownership of downloads.

2. Detecting File Size When Content-Length Is Missing in Requests

You’re checking the request headers for Content-Length, but that’s the size of the data sent to the server—not the file you’re downloading. The file size lives in the response headers sent by the server. Plus, some servers use chunked transfer encoding (no Content-Length), so we need a fallback approach.

Solution: Use HEAD Requests to Pre-Check File Size

Before allowing the download request to proceed, send a HEAD request to the server to fetch the file size without downloading the entire file. If it exceeds your limit, abort the original request.

await page.SetRequestInterceptionAsync(true);
page.Request += async (sender, e) => 
{
    var request = e.Request;
    // Target only downloadable resources (adjust filters to match your use case)
    var isDownloadable = request.ResourceType == ResourceType.Other 
                        || request.Url.EndsWith(".zip") 
                        || request.Url.EndsWith(".exe") 
                        || request.Url.EndsWith(".pdf");

    if (isDownloadable)
    {
        try
        {
            // Send a HEAD request to get file size without full download
            using var httpClient = new HttpClient();
            var headRequest = new HttpRequestMessage(HttpMethod.Head, request.Url);
            
            // Copy original request headers (e.g., auth cookies) to access restricted resources
            foreach (var header in request.Headers)
            {
                if (!headRequest.Headers.TryAddWithoutValidation(header.Key, header.Value))
                {
                    headRequest.Content?.Headers.TryAddWithoutValidation(header.Key, header.Value);
                }
            }

            var headResponse = await httpClient.SendAsync(headRequest);
            if (headResponse.IsSuccessStatusCode)
            {
                // Check if Content-Length is available
                if (headResponse.Content.Headers.ContentLength.HasValue)
                {
                    var fileSize = headResponse.Content.Headers.ContentLength.Value;
                    if (fileSize > GIVENSIZE)
                    {
                        await request.AbortAsync();
                        Console.WriteLine($"Aborted download: {request.Url} (size exceeds limit)");
                        return;
                    }
                }
                else
                {
                    // Handle chunked transfer (no Content-Length)
                    // Option 1: Allow request and monitor download progress (more complex)
                    // Option 2: Abort if size can't be verified
                    // await request.AbortAsync();
                    // Console.WriteLine($"Aborted download: {request.Url} (no Content-Length available)");
                    // return;
                }
            }
        }
        catch (Exception ex)
        {
            // Handle cases where HEAD requests aren't supported by the server
            Console.WriteLine($"Failed to check file size: {ex.Message}");
        }
    }

    // Allow request to proceed if all checks pass
    await request.ContinueAsync();
};

Key Notes:

  • Copying headers from the original request ensures you can access authenticated or restricted resources.
  • For servers that don’t support HEAD requests, adjust the logic to either allow the download (and monitor progress if needed) or abort it based on your requirements.
  • If you need to handle chunked transfers, you’d need to listen to the Response event and track downloaded bytes, then terminate the request if it exceeds your limit—this is more involved but covers edge cases.

内容的提问来源于stack exchange,提问作者Inside Man

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:05:46