You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

for循环中使用Goroutine引发异常行为的问题排查

Go语言爬虫练习问题解答

问题背景

我在完成《Go语言之旅》的Web Crawler练习时,基于一个并发Mutex实现修改适配预定义签名,但爬虫在URL树的第二层就停止了。调试时发现打印语句的位置不同会导致奇怪的输出:

情况1:打印语句在Goroutine外部

var done sync.WaitGroup
for _, u := range urls {
    done.Add(1)
    fmt.Printf("enter: %s\n", u) // 打印在外部
    go func(url string) {
        defer done.Done()
        Crawl(u, depth-1, fetcher, f)
    }(u)
}
done.Wait()

输出符合预期,但爬虫会停止:

enter: https://golang.org/pkg/
enter: https://golang.org/cmd/

情况2:打印语句在Goroutine内部

var done sync.WaitGroup
for _, u := range urls {
    done.Add(1)
    go func(url string) {
        defer done.Done()
        fmt.Printf("enter: %s\n", u) // 打印在内部
        Crawl(u, depth-1, fetcher, f)
    }(u)
}
done.Wait()

输出重复:

enter: https://golang.org/cmd/
enter: https://golang.org/cmd/

我有两个问题:

  1. 第二种情况中,为什么enter: https://golang.org/cmd/会被打印两次?
  2. 为什么Crawl函数遇到错误就停止,而非继续遍历URL树?

注:第二个问题可能与第一个相关,我故意在Goroutine内使用u而非url来复现该bug。

完整代码如下:

package main

import (
    "fmt"
    "sync"
)

type Fetcher interface {
    // Fetch returns the body of URL and
    // a slice of URLs found on that page.
    Fetch(url string) (body string, urls []string, err error)
}

type fetchState struct {
    mu sync.Mutex
    fetched map[string]bool
}

// Crawl uses fetcher to recursively crawl
// pages starting with url, to a maximum of depth.
func Crawl(url string, depth int, fetcher Fetcher, f *fetchState) {
    // TODO: Fetch URLs in parallel.
    // TODO: Don't fetch the same URL twice.
    // This implementation doesn't do either:
    f.mu.Lock()
    already := f.fetched[url]
    f.fetched[url] = true
    f.mu.Unlock()
    
    if already {
        return
    }
    
    if depth <= 0 {
        return
    }
    body, urls, err := fetcher.Fetch(url)
    if err != nil {
        fmt.Println(err)
        return
    }
    fmt.Printf("found: %s %q\n", url, body)
    
    var done sync.WaitGroup
    for _, u := range urls {
        done.Add(1)
        go func(url string) {
            defer done.Done()
            fmt.Printf("enter: %s\n", u)
            Crawl(u, depth-1, fetcher, f)
        }(u)
    }
    done.Wait()
    
    return
}

func makeState() *fetchState{
    f := &fetchState{}
    f.fetched = make(map[string]bool)
    return f
}

func main() {
    Crawl("https://golang.org/", 4, fetcher, makeState())
}

// fakeFetcher is Fetcher that returns canned results.
type fakeFetcher map[string]*fakeResult

type fakeResult struct {
    body string
    urls []string
}

func (f fakeFetcher) Fetch(url string) (string, []string, error) {
    if res, ok := f[url]; ok {
        return res.body, res.urls, nil
    }
    return "", nil, fmt.Errorf("not found: %s", url)
}

// fetcher is a populated fakeFetcher.
var fetcher = fakeFetcher{
    "https://golang.org/": &fakeResult{
        "The Go Programming Language",
        []string{
            "https://golang.org/pkg/",
            "https://golang.org/cmd/",
        },
    },
    "https://golang.org/pkg/": &fakeResult{
        "Packages",
        []string{
            "https://golang.org/",
            "https://golang.org/cmd/",
            "https://golang.org/pkg/fmt/",
            "https://golang.org/pkg/os/",
        },
    },
    "https://golang.org/pkg/fmt/": &fakeResult{
        "Package fmt",
        []string{
            "https://golang.org/",
            "https://golang.org/pkg/",
        },
    },
    "https://golang.org/pkg/os/": &fakeResult{
        "Package os",
        []string{
            "https://golang.org/",
            "https://golang.org/pkg/",
        },
    },
}

问题解答

1. 为什么第二种情况会重复打印同一个URL?

这是Go语言中循环变量捕获的经典坑:

  • for循环中的变量u是同一个内存地址的变量,循环迭代时会不断更新它的值。
  • Goroutine启动后不会立即执行,当它们真正开始运行时,循环可能已经完成所有迭代,此时u的值已经是循环的最后一个元素(即https://golang.org/cmd/)。
  • 你在匿名函数中使用的是外部的u而非传入的参数url,导致两个Goroutine都读取到u的最终值,因此重复打印同一个URL。

修复方法是在Goroutine内部使用传入的参数url,而非外部循环变量u:

go func(url string) {
    defer done.Done()
    fmt.Printf("enter: %s\n", url) // 使用参数url而非外部u
    Crawl(url, depth-1, fetcher, f)
}(u)

2. 为什么爬虫会停止遍历?

这个问题直接由第一个问题的bug导致:

  • 由于Goroutine错误地使用循环变量u,两个Goroutine都去爬取https://golang.org/cmd/。
  • 第一个Goroutine爬取时,会将fetched["https://golang.org/cmd/"]标记为true;第二个Goroutine进入Crawl函数后,检查到该URL已被爬取,直接返回,没有做任何后续操作。
  • 原本应该被爬取的https://golang.org/pkg/完全没有被处理,导致爬虫看起来在第二层就停止了。

另外需要澄清:Crawl函数中Fetch出错时的return只会终止当前Goroutine,不会影响其他并发执行的Goroutine。你看到的"停止"并不是因为错误,而是因为循环变量捕获的bug导致本该被爬取的URL没有被处理。


内容的提问来源于stack exchange,提问作者Liao-Liao Liu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 01:48:38