C# 使用Selenium(Chrome)开发ASP.NET爬虫如何获取下载完成的文件名
Selenium C# 批量下载并实时获取已完成文件名实现方案
前置配置:浏览器参数设置
首先需要在WebDriver初始化时配置下载规则,避免弹出系统下载弹窗,同时指定固定的下载目录用于后续监控:
var downloadPath = @"C:\Your\Custom\Download\Path"; var options = new ChromeOptions(); // 配置默认下载目录 options.AddUserProfilePreference("download.default_directory", downloadPath); // 禁用下载弹窗 options.AddUserProfilePreference("download.prompt_for_download", false); options.AddUserProfilePreference("download.directory_upgrade", true); // 禁用安全下载拦截,避免ASP.NET站点返回的文件被浏览器误拦截 options.AddUserProfilePreference("safebrowsing.enabled", false); // 初始化Driver var _driver = new ChromeDriver(options);
核心方法:下载状态监控
封装方法实时监控下载目录,判断新文件是否下载完成,返回最终文件名:
/// <summary> /// 等待下载完成并返回最新下载的文件完整路径 /// </summary> /// <param name="downloadDir">下载目录路径</param> /// <param name="snapshotFiles">下载触发前目录内的已有文件快照</param> /// <param name="timeoutSeconds">超时时间,默认30秒</param> /// <returns>下载完成的文件完整路径</returns> /// <exception cref="TimeoutException">下载超时时抛出</exception> public string WaitForDownloadComplete(string downloadDir, HashSet<string> snapshotFiles, int timeoutSeconds = 30) { var endTime = DateTime.Now.AddSeconds(timeoutSeconds); while (DateTime.Now < endTime) { // 获取当前目录所有文件 var currentFiles = Directory.GetFiles(downloadDir).ToHashSet(); // 筛选本次触发下载后新增的文件 var newFiles = currentFiles.Except(snapshotFiles).ToList(); if (newFiles.Any()) { // 排除未下载完成的临时文件,Chrome临时后缀为.crdownload、Edge为.part,可根据使用的浏览器调整规则 var completedFile = newFiles.FirstOrDefault(f => !f.EndsWith(".crdownload") && !f.EndsWith(".tmp") && !f.EndsWith(".part")); if (completedFile != null) { return completedFile; } } // 间隔100ms轮询,避免占用过高CPU资源 Thread.Sleep(100); } throw new TimeoutException("文件下载超时"); }
批量下载逻辑改造
把原有触发下载代码和监控逻辑结合,实现逐个下载并实时获取文件名:
var downloadPath = @"C:\Your\Custom\Download\Path"; // 替换为实际需要下载的文件总数 int totalCount = 10; for (int i = 1; i <= totalCount; i++) { // 1. 下载前先获取当前目录的文件快照 var beforeSnapshot = Directory.GetFiles(downloadPath).ToHashSet(); // 2. 触发下载(原有逻辑) IWebElement tableElementaa = _driver.FindElement(By.Id($"cph_gvS02_btnDown_{i - 1}")); tableElementaa.Click(); // 3. 等待下载完成,获取文件路径 var downloadedFilePath = WaitForDownloadComplete(downloadPath, beforeSnapshot); // 此处已拿到下载完成的文件路径,可直接做重命名、移动等后续处理 Console.WriteLine($"第{i}个文件下载完成:{Path.GetFileName(downloadedFilePath)}"); }
注意事项
- 批量下载建议逐个触发,避免同时下载多个文件导致新增文件识别混乱;如果需要并发下载,可以为每个下载任务创建独立的子下载目录,分别监控即可
- 超时时间可根据你要下载的文件大小调整,避免大文件未下载完成就触发超时
- 如果使用Firefox/Edge等其他浏览器,只需要调整对应浏览器的下载配置参数,以及临时文件的后缀判断规则即可
内容的提问来源于stack exchange,提问作者Lawliet.Lin
相关产品推荐
相关产品推荐

