for循环中使用Goroutine引发异常行为的问题排查
Go语言爬虫练习问题解答
问题背景
我在完成《Go语言之旅》的Web Crawler练习时,基于一个并发Mutex实现修改适配预定义签名,但爬虫在URL树的第二层就停止了。调试时发现打印语句的位置不同会导致奇怪的输出:
情况1:打印语句在Goroutine外部
var done sync.WaitGroup for _, u := range urls { done.Add(1) fmt.Printf("enter: %s\n", u) // 打印在外部 go func(url string) { defer done.Done() Crawl(u, depth-1, fetcher, f) }(u) } done.Wait()
输出符合预期,但爬虫会停止:
enter: https://golang.org/pkg/ enter: https://golang.org/cmd/
情况2:打印语句在Goroutine内部
var done sync.WaitGroup for _, u := range urls { done.Add(1) go func(url string) { defer done.Done() fmt.Printf("enter: %s\n", u) // 打印在内部 Crawl(u, depth-1, fetcher, f) }(u) } done.Wait()
输出重复:
enter: https://golang.org/cmd/ enter: https://golang.org/cmd/
我有两个问题:
- 第二种情况中,为什么
enter: https://golang.org/cmd/会被打印两次? - 为什么Crawl函数遇到错误就停止,而非继续遍历URL树?
注:第二个问题可能与第一个相关,我故意在Goroutine内使用u而非url来复现该bug。
完整代码如下:
package main import ( "fmt" "sync" ) type Fetcher interface { // Fetch returns the body of URL and // a slice of URLs found on that page. Fetch(url string) (body string, urls []string, err error) } type fetchState struct { mu sync.Mutex fetched map[string]bool } // Crawl uses fetcher to recursively crawl // pages starting with url, to a maximum of depth. func Crawl(url string, depth int, fetcher Fetcher, f *fetchState) { // TODO: Fetch URLs in parallel. // TODO: Don't fetch the same URL twice. // This implementation doesn't do either: f.mu.Lock() already := f.fetched[url] f.fetched[url] = true f.mu.Unlock() if already { return } if depth <= 0 { return } body, urls, err := fetcher.Fetch(url) if err != nil { fmt.Println(err) return } fmt.Printf("found: %s %q\n", url, body) var done sync.WaitGroup for _, u := range urls { done.Add(1) go func(url string) { defer done.Done() fmt.Printf("enter: %s\n", u) Crawl(u, depth-1, fetcher, f) }(u) } done.Wait() return } func makeState() *fetchState{ f := &fetchState{} f.fetched = make(map[string]bool) return f } func main() { Crawl("https://golang.org/", 4, fetcher, makeState()) } // fakeFetcher is Fetcher that returns canned results. type fakeFetcher map[string]*fakeResult type fakeResult struct { body string urls []string } func (f fakeFetcher) Fetch(url string) (string, []string, error) { if res, ok := f[url]; ok { return res.body, res.urls, nil } return "", nil, fmt.Errorf("not found: %s", url) } // fetcher is a populated fakeFetcher. var fetcher = fakeFetcher{ "https://golang.org/": &fakeResult{ "The Go Programming Language", []string{ "https://golang.org/pkg/", "https://golang.org/cmd/", }, }, "https://golang.org/pkg/": &fakeResult{ "Packages", []string{ "https://golang.org/", "https://golang.org/cmd/", "https://golang.org/pkg/fmt/", "https://golang.org/pkg/os/", }, }, "https://golang.org/pkg/fmt/": &fakeResult{ "Package fmt", []string{ "https://golang.org/", "https://golang.org/pkg/", }, }, "https://golang.org/pkg/os/": &fakeResult{ "Package os", []string{ "https://golang.org/", "https://golang.org/pkg/", }, }, }
问题解答
1. 为什么第二种情况会重复打印同一个URL?
这是Go语言中循环变量捕获的经典坑:
- for循环中的变量
u是同一个内存地址的变量,循环迭代时会不断更新它的值。 - Goroutine启动后不会立即执行,当它们真正开始运行时,循环可能已经完成所有迭代,此时
u的值已经是循环的最后一个元素(即https://golang.org/cmd/)。 - 你在匿名函数中使用的是外部的
u而非传入的参数url,导致两个Goroutine都读取到u的最终值,因此重复打印同一个URL。
修复方法是在Goroutine内部使用传入的参数url,而非外部循环变量u:
go func(url string) { defer done.Done() fmt.Printf("enter: %s\n", url) // 使用参数url而非外部u Crawl(url, depth-1, fetcher, f) }(u)
2. 为什么爬虫会停止遍历?
这个问题直接由第一个问题的bug导致:
- 由于Goroutine错误地使用循环变量
u,两个Goroutine都去爬取https://golang.org/cmd/。 - 第一个Goroutine爬取时,会将
fetched["https://golang.org/cmd/"]标记为true;第二个Goroutine进入Crawl函数后,检查到该URL已被爬取,直接返回,没有做任何后续操作。 - 原本应该被爬取的
https://golang.org/pkg/完全没有被处理,导致爬虫看起来在第二层就停止了。
另外需要澄清:Crawl函数中Fetch出错时的return只会终止当前Goroutine,不会影响其他并发执行的Goroutine。你看到的"停止"并不是因为错误,而是因为循环变量捕获的bug导致本该被爬取的URL没有被处理。
内容的提问来源于stack exchange,提问作者Liao-Liao Liu
相关产品推荐
相关产品推荐

