You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

.Net 6控制台应用使用HttpClient下载资源时间歇性404问题排查

间歇性404问题排查:代码还是网络原因?

问题背景

在.Net 6控制台应用中爬取某无法后台控制的网页,提取链接后通过HttpClient将网页内容下载到本地文件,随后遍历链接再次下载对应文件。在家用网络运行时会间歇性出现404服务器错误,但在网速更快的咖啡店网络运行时无失败情况。所有链接在Chrome和Postman中均可正常访问,已尝试不同下载方式、URL及网络连接,未找到解决方案,需确认问题根源及解决办法。

现有代码

Downloader类实现

public class Downloader
{
    /// <summary>
    /// The Downloader class's HttpClient instance
    /// </summary>
    private HttpClient client { get; }

    /// <summary>
    /// Downloader class that provides methods for downloading web pages and files from internet based resources.
    /// </summary>
    public Downloader()
    {   
        this.client = new HttpClient();
    }

    /// <summary>
    /// Function to assist in downloading web pages and files to local memory
    /// </summary>
    /// <returns>
    /// A string containing the file name of the downloaded web asset
    /// </returns>
    public async Task<string> GetWebAsset(string url)
    {
        // File name is the end of the url or temp.txt
        string fileName = Path.GetFileName(url) != string.Empty ? Path.GetFileName(url) : "temp.txt";

        // Paths can contain ?ver=xyz after their extension, remove additional params
        if (fileName.Contains("?")) {
            string[] fileNameParts = fileName.Split("?");
            fileName = fileNameParts[0];
        }

        string newFileName = fileName;

        // Make sure the current file name doesn't exist in memory, append a number to it if it does
        int i = 0;
        while(File.Exists(newFileName))
        {
            newFileName = fileName;
            i++;
            string[] fileNameParts = fileName.Split(".");
            newFileName = fileNameParts[0] + i.ToString() + "." + fileNameParts[1];
        }

        fileName = newFileName;

        using (HttpResponseMessage response = await this.client.GetAsync(url, HttpCompletionOption.ResponseHeadersRead))
        using (Stream streamToReadFrom = await response.Content.ReadAsStreamAsync())
        {
            using (Stream fileToWriteTo = File.Open(fileName, FileMode.CreateNew))
            {
                await streamToReadFrom.CopyToAsync(fileToWriteTo);
            }
        }

        // using var downloadStream = await this.client.GetStreamAsync(url);
        // using var fileStream = new FileStream(fileName, FileMode.CreateNew);

        // await downloadStream.CopyToAsync(fileStream);

        return fileName;
    }
}

调用示例

Downloader dl = new Downloader();
string google = await dl.GetWebAsset("https://www.google.com/");

间歇性错误内容

当出现404时,错误页面被保存为文件(如MANAGEMENT.zip),内容如下:

<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.0 Strict//EN" "http://www.w3.org/TR/xhtml1/DTD/xhtml1-strict.dtd">
<html xmlns="http://www.w3.org/1999/xhtml">
<head>
<meta http-equiv="Content-Type" content="text/html; charset=iso-8859-1"/>
<title>404 - File or directory not found.</title>
<style type="text/css">
<!--
body{margin:0;font-size:.7em;font-family:Verdana, Arial, Helvetica, sans-serif;background:#EEEEEE;}
fieldset{padding:0 15px 10px 15px;} 
h1{font-size:2.4em;margin:0;color:#FFF;}
h2{font-size:1.7em;margin:0;color:#CC0000;} 
h3{font-size:1.2em;margin:10px 0 0 0;color:#000000;} 
#header{width:96%;margin:0 0 0 0;padding:6px 2% 6px 2%;font-family:"trebuchet MS", Verdana, sans-serif;color:#FFF;
background-color:#555555;}
#content{margin:0 0 0 2%;position:relative;}
.content-container{background:#FFF;width:96%;margin-top:8px;padding:10px;position:relative;}
-->
</style>
</head>
<body>
<div id="header"><h1>Server Error</h1></div>
<div id="content">
 <div class="content-container"><fieldset>
  <h2>404 - File or directory not found.</h2>
  <h3>The resource you are looking for might have been removed, had its name changed, or is temporarily unavailable.</h3>
 </fieldset></div>
</div>
</body>
</html>

问题分析

网络因素可能性

  1. 家用网络稳定性:家用网络带宽波动大、延迟高,甚至存在丢包情况,导致请求无法完整到达服务器,或服务器响应超时后返回404;而咖啡店网络带宽稳定,请求能正常交互。
  2. DNS解析问题:家用网络DNS服务器可能存在解析延迟或偶尔解析错误,导致请求发送到错误的服务器节点,返回404。
  3. 运营商/NAT限制:家用网络多为动态IP,运营商的NAT转发可能导致请求被服务器判定为异常流量,触发拦截返回404;咖啡店网络可能使用公网IP或更宽松的NAT策略。

代码层面问题

  1. HttpClient实例管理不当:每次创建Downloader都实例化新的HttpClient,会快速耗尽HTTP连接池,导致请求排队超时,进而引发异常或被服务器拒绝。
  2. 未处理HTTP错误状态码:代码中未检查response.IsSuccessStatusCode,即使服务器返回404,仍会将错误页面保存为目标文件,无法及时发现请求失败。
  3. 文件名处理逻辑缺陷:当文件名没有后缀时(如test),Split(".")会得到长度为1的数组,访问fileNameParts[1]会抛出索引越界异常,影响后续请求。
  4. 缺少请求头模拟:未设置User-Agent等浏览器常用请求头,部分服务器会将此类请求判定为爬虫,返回404或拦截。
  5. 无重试机制:家用网络间歇性故障时,请求失败后没有重试逻辑,直接导致下载失败。

解决办法

网络方面优化

  • 测试家用网络连通性:使用ping、tracert命令测试目标服务器的延迟和丢包率,确认网络稳定性。
  • 更换DNS服务器:将家用网络DNS改为公共DNS(如8.8.8.8、114.114.114.114),减少解析错误概率。
  • 检查防火墙/代理:确认家用网络没有防火墙、路由器规则或代理拦截目标网站的请求。

代码层面优化

1. 改用单例HttpClient或IHttpClientFactory

避免频繁创建HttpClient,复用实例以节省连接池资源:

public class Downloader
{
    // 使用静态单例HttpClient
    private static readonly HttpClient _client = new HttpClient();

    public Downloader() { }

    // 后续方法复用_client
}

或使用IHttpClientFactory(推荐,适合依赖注入场景)。

2. 添加HTTP错误处理

在获取响应后检查状态码,及时抛出异常:

using (HttpResponseMessage response = await _client.GetAsync(url, HttpCompletionOption.ResponseHeadersRead))
{
    // 检查状态码,非成功则抛出异常
    response.EnsureSuccessStatusCode();
    using (Stream streamToReadFrom = await response.Content.ReadAsStreamAsync())
    {
        using (Stream fileToWriteTo = File.Open(fileName, FileMode.CreateNew))
        {
            await streamToReadFrom.CopyToAsync(fileToWriteTo);
        }
    }
}

3. 完善文件名处理逻辑

处理无后缀的文件名,避免索引越界:

string newFileName = fileName;
int i = 0;
while (File.Exists(newFileName))
{
    string extension = Path.GetExtension(fileName);
    string fileNameWithoutExt = Path.GetFileNameWithoutExtension(fileName);
    newFileName = $"{fileNameWithoutExt}{i}{extension}";
    i++;
}
fileName = newFileName;

4. 添加浏览器请求头

模拟浏览器请求,避免被服务器拦截:

private static readonly HttpClient _client = new HttpClient();

static Downloader()
{
    _client.DefaultRequestHeaders.UserAgent.ParseAdd("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36");
    _client.DefaultRequestHeaders.Accept.ParseAdd("text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8");
}

5. 实现重试机制

使用Polly库添加重试策略,处理间歇性网络错误:
首先安装Polly包:Install-Package Polly
然后添加重试逻辑:

// 定义重试策略:针对HttpRequestException、404等状态码重试3次,每次间隔1秒
var retryPolicy = Policy
    .Handle<HttpRequestException>()
    .OrResult<HttpResponseMessage>(r => r.StatusCode == HttpStatusCode.NotFound)
    .WaitAndRetryAsync(3, retryAttempt => TimeSpan.FromSeconds(1));

// 在请求中使用策略
using (HttpResponseMessage response = await retryPolicy.ExecuteAsync(() => 
    _client.GetAsync(url, HttpCompletionOption.ResponseHeadersRead)))
{
    response.EnsureSuccessStatusCode();
    // 后续下载逻辑
}

内容的提问来源于stack exchange,提问作者hudsonsc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 18:45:47