Go语言如何检测URL是否为可下载的文件资源?
实现方案
核心思路是提前校验HTTP响应头特征,判断目标是否为文件资源,校验不通过直接返回错误,避免生成无效本地文件,常用实现方案如下:
方案1:先发HEAD请求预校验(推荐,节省带宽)
HEAD请求只会获取响应头,不会下载响应体,适合大文件场景的预校验,校验通过后再发GET请求下载即可。
需要校验的核心响应头字段:
Content-Disposition:正常文件下载响应一般会携带attachment标识,同时会包含文件名信息Content-Type:过滤非文件类MIME类型,比如text/html大概率是网页、接口返回等非文件资源,可以根据业务需求设置允许的MIME白名单Content-Length:正常静态文件一般会携带明确的内容长度,如果没有该字段且非分片传输的场景,大概率是动态生成的非静态文件资源
修改后代码示例:
package main import ( "errors" "fmt" "io" "net/http" "os" "strings" ) // 允许的文件MIME类型白名单,可根据业务需求调整 var allowMimeTypes = map[string]bool{ "text/plain": true, "application/pdf": true, "image/jpeg": true, "image/png": true, "application/zip": true, "application/octet-stream": true, // 通用二进制文件流 } func main() { fileUrl := "http://example.com/file.txt" err := DownloadFile("./example.txt", fileUrl) if err != nil { panic(err) } fmt.Println("Downloaded: " + fileUrl) } // DownloadFile 下载url到本地文件,下载前会校验是否为文件资源 func DownloadFile(filepath string, url string) error { // 第一步:先发HEAD请求预校验 headResp, err := http.Head(url) if err != nil { return fmt.Errorf("HEAD请求失败: %w", err) } defer headResp.Body.Close() if headResp.StatusCode < 200 || headResp.StatusCode >= 300 { return errors.New("请求资源不存在或无访问权限") } // 校验Content-Disposition是否为附件 contentDisposition := headResp.Header.Get("Content-Disposition") if contentDisposition != "" && !strings.Contains(contentDisposition, "attachment") { return errors.New("目标资源不是可下载的文件") } // 校验MIME类型是否在白名单内 contentType := headResp.Header.Get("Content-Type") // 去掉charset等后缀 if idx := strings.Index(contentType, ";"); idx != -1 { contentType = contentType[:idx] } if !allowMimeTypes[contentType] { return fmt.Errorf("不允许下载的文件类型: %s", contentType) } // 校验通过,发起GET请求下载 resp, err := http.Get(url) if err != nil { return err } defer resp.Body.Close() if resp.StatusCode < 200 || resp.StatusCode >= 300 { return errors.New("下载请求失败") } // 创建本地文件 out, err := os.Create(filepath) if err != nil { return err } defer out.Close() // 写入内容 _, err = io.Copy(out, resp.Body) return err }
方案2:GET请求后先校验响应头再写入文件
如果目标服务器禁止HEAD请求(返回405 Method Not Allowed),可以直接发起GET请求,拿到响应后先校验头信息,校验不通过直接返回错误,不创建本地文件即可,核心逻辑如下:
func DownloadFile(filepath string, url string) error { resp, err := http.Get(url) if err != nil { return err } defer resp.Body.Close() if resp.StatusCode < 200 || resp.StatusCode >= 300 { return errors.New("请求资源不存在或无访问权限") } // 校验逻辑和HEAD请求校验完全一致 contentDisposition := resp.Header.Get("Content-Disposition") if contentDisposition != "" && !strings.Contains(contentDisposition, "attachment") { return errors.New("目标资源不是可下载的文件") } contentType := resp.Header.Get("Content-Type") if idx := strings.Index(contentType, ";"); idx != -1 { contentType = contentType[:idx] } if !allowMimeTypes[contentType] { return fmt.Errorf("不允许下载的文件类型: %s", contentType) } // 校验通过再创建本地文件 out, err := os.Create(filepath) if err != nil { return err } defer out.Close() _, err = io.Copy(out, resp.Body) return err }
注意事项
- MIME白名单可以根据业务场景灵活调整,如果不需要限制文件类型,也可以只保留
Content-Disposition和状态码的校验 - 遇到分片下载(
Transfer-Encoding: chunked)的场景,Content-Length字段会为空,这种情况如果确认是合法文件资源,可以跳过Content-Length的校验
内容的提问来源于stack exchange,提问作者meena sushanth
相关产品推荐
相关产品推荐

