Colly Go包:如何判断Timeout错误并仅在超时后重试?
解决Colly爬虫仅在Timeout错误时重试的问题
Colly没有直接暴露可直接比对的超时错误常量,要判断请求是否因超时失败,需借助底层net包的错误类型来识别:
- 添加必要导入:在代码的import块中加入
"errors"和"net"两个包,用于错误类型判断。 - 修改错误判断逻辑:在
OnError回调中,将错误匹配为net.Error类型,并检查该错误的Timeout()方法返回值,以此确认是超时错误。 - 优化冗余请求:原代码中
c.Request和c.Visit重复调用会导致同一链接被请求两次,建议只保留其中一个。
修改后的完整代码
package main import ( "crypto/tls" "errors" "fmt" "github.com/gocolly/colly" "net" "net/http" "os" "strings" ) func main() { crawl() } func crawl() { httpMethod := "https" domains := []string{ "www.myweb1.com", "www.myweb2.com", } for _, domain := range domains { // Instantiate default collector c := colly.NewCollector( colly.Async(true), colly.AllowedDomains(domain), ) c.Limit(&colly.LimitRule{Parallelism: 100}) /* default = 10s = 1000000000 nanoseconds: */ c.SetRequestTimeout(10 * 1e9) c.WithTransport(&http.Transport{ TLSClientConfig:&tls.Config{InsecureSkipVerify: true}, }) // On every a element which has href attribute call callback c.OnHTML("a[href]", func(e *colly.HTMLElement) { url := e.Request.URL.String() link := e.Request.AbsoluteURL(e.Attr("href")) // Only those links are visited which are in AllowedDomains // create a new context to remember the referer (in case of error) for _, domain_ok := range domains { if strings.Contains(link, domain_ok) { ctx := colly.NewContext() ctx.Put("Referrer", url) // 只保留c.Request即可,避免重复请求 c.Request(http.MethodGet, link, nil, ctx, nil) break } } if !strings.Contains(link, domain) { fmt.Fprintf(os.Stdout, "Ignoring %s (%s)\n", link, url) } }) // Before making a request print "Visiting ..." c.OnRequest(func(r *colly.Request) { fmt.Println("Visiting", r.URL.String()) }) c.OnError(func(resp *colly.Response, err error) { url := resp.Request.URL.String() fmt.Fprintf( os.Stdout, "ERR on URL: %s (from: %s), error: %s\n", url, resp.Request.Ctx.Get("Referrer"), err, ) var netErr net.Error // 判断是否为超时错误 if errors.As(err, &netErr) && netErr.Timeout() { fmt.Fprintf(os.Stdout, "Retry: '%s'\n", url) resp.Request.Retry() } }) urlBase := fmt.Sprintf("%s://%s", httpMethod, domain) fmt.Println("Scraping: ", urlBase) c.Visit(urlBase) c.Wait() } }
关键说明
- 使用
errors.As而非直接类型断言,能更好地处理错误链中的嵌套错误,确保不会遗漏包装后的超时错误。 resp.Request.Retry()会使用相同的上下文和请求参数重新发起请求,符合重试需求。
内容的提问来源于stack exchange,提问作者Olivier Pons
相关产品推荐
相关产品推荐

