You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Go网络爬虫并发控制异常:goroutine数量超预期三倍,寻求排查方案

问题分析与解决方案

首先咱们拆解下核心问题:你看到的并发数是设置值3倍的根源,在于**runtime.NumGoroutine()统计的是所有goroutine——包括你自己启动的,还有Go标准库HTTP客户端创建的底层goroutine**。

1. 为什么并发数是设置值的3倍?

你的信号量逻辑其实是有效的,它确实把你主动启动的goroutine数量限制在了80个。但每个http.Get()请求背后,标准库的http.Transport会自动创建额外的goroutine:

  • 一个负责建立TCP连接或从连接池获取连接
  • 另一个负责读取HTTP响应内容
    这两个额外的goroutine会被runtime.NumGoroutine()统计进去,所以80个自定义goroutine + 160个标准库goroutine,正好是240左右,和你观察到的比例完全匹配。

2. 你的代码里还有两个关键隐患

(1)切片操作的线程不安全问题

你在多个goroutine里直接调用scrapeResults = append(scrapeResults, ...),这是严重的线程不安全行为。切片的底层数组是共享的,多个goroutine同时append会导致数据竞争,可能出现结果丢失、重复甚至程序panic。

(2)HTTP客户端默认行为导致额外goroutine

默认的http.Client使用的Transport参数比较宽松,允许大量闲置连接和并发连接,从而催生更多底层goroutine。

3. 解决方案与排查步骤

(1)准确统计自定义goroutine数量

如果你只想看自己启动的goroutine数量,别用runtime.NumGoroutine(),改用原子计数器:

import "sync/atomic"

func (B BrandScraper) ScrapeUrls(URLs ...string) []scrapeResponse {
    concurrent := 80
    semaphoreChan := make(chan struct{}, concurrent)
    var activeGoroutines int32 // 原子计数器,保证线程安全

    for _, URL := range URLs {
        semaphoreChan <- struct{}{}
        go func(URL string) {
            atomic.AddInt32(&activeGoroutines, 1)
            defer func() {
                <-semaphoreChan
                atomic.AddInt32(&activeGoroutines, -1)
            }()
            // 打印仅统计自定义goroutine的数量
            fmt.Printf("#custom goroutines: %d\n", atomic.LoadInt32(&activeGoroutines))
            // ... 其他逻辑 ...
        }(URL)
    }
    // ... 结果收集逻辑 ...
}

这样你会看到自定义goroutine数量稳定在80左右,符合预期。

(2)修复线程不安全的结果收集

改用带缓冲的通道接收结果,避免多个goroutine同时操作切片:

func (B BrandScraper) ScrapeUrls(URLs ...string) []scrapeResponse {
    concurrent := 80
    semaphoreChan := make(chan struct{}, concurrent)
    resultsChan := make(chan scrapeResponse, len(URLs)) // 缓冲大小等于URL总数

    for _, URL := range URLs {
        semaphoreChan <- struct{}{}
        go func(URL string) {
            defer func() { <-semaphoreChan }()
            resultsChan <- B.getIndividualScrape(URL)
        }(URL)
    }

    // 安全收集所有结果
    scrapeResults := make([]scrapeResponse, 0, len(URLs))
    for i := 0; i < len(URLs); i++ {
        scrapeResults = append(scrapeResults, <-resultsChan)
    }
    close(resultsChan)

    return scrapeResults
}

(3)控制HTTP客户端的额外goroutine

自定义http.Transport参数,限制并发连接数,从而减少底层goroutine的数量:

import "net/http"
import "time"

func (B BrandScraper) doGetRequest(URL string) io.Reader {
    // 自定义Transport,严格限制连接数
    transport := &http.Transport{
        MaxConnsPerHost:     80,    // 每个主机的最大并发连接数
        MaxIdleConns:        80,    // 全局最大闲置连接数
        MaxIdleConnsPerHost: 80,    // 每个主机的最大闲置连接数
        IdleConnTimeout:     30 * time.Second, // 闲置连接超时回收
    }
    client := &http.Client{Transport: transport}

    resp, err := client.Get(URL)
    if err != nil {
        log.Fatal(err)
    }
    defer resp.Body.Close() // 把Close放在defer里,确保无论读取成功与否都关闭连接

    body, _ := ioutil.ReadAll(resp.Body)
    return bytes.NewReader(body)
}

(4)排查goroutine泄漏的方法

如果怀疑有goroutine泄漏,用Go的pprof工具深度分析:

  1. 在代码中导入net/http/pprof,并启动调试服务器:
    import _ "net/http/pprof"
    
    func main() {
        // 启动pprof调试服务
        go func() {
            log.Println(http.ListenAndServe("localhost:6060", nil))
        }()
        // 你的爬虫逻辑...
    }
    
  2. 运行程序后,访问http://localhost:6060/debug/pprof/goroutine?debug=2,查看所有goroutine的栈信息。如果存在泄漏,你会看到长时间阻塞的goroutine(比如等待未关闭的通道、未释放的网络连接等)。

内容的提问来源于stack exchange,提问作者dircrys

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 09:16:20