使用Go-Colly爬取网站返回空切片问题排查求助
Colly爬虫返回空Product切片问题排查
问题描述
使用Go语言的colly库爬取网站时,OnHTML回调函数中能正常输出product变量,但最终返回的Product切片为空。怀疑是切片追加操作有误,附上代码请求排查。
scraper.go 原代码
package scraper import ( "fmt" "strings" "github.com/gocolly/colly" "github.com/gocolly/colly/extensions" ) type Product struct { name string fullPrice string url string } func Scraper(site string) []Product { products := []Product{} c := colly.NewCollector() replacer := strings.NewReplacer("R$", "", ",", ".") c.OnHTML("div#column-main-content", func(e *colly.HTMLElement) { fullPrice := e.ChildText("span.m7nrfa-0.eJCbzj.sc-ifAKCX.ANnoQ") product := Product{ name: e.ChildText("h2"), fullPrice: replacer.Replace(fullPrice), url: e.ChildAttr("a.sc-1fcmfeb-2.iezWpY", "href"), } fmt.Println(product) products = append(products, product) }) fmt.Println(products) c.OnRequest(func(r *colly.Request) { fmt.Println("Visiting", r.URL) }) c.OnError(func(r *colly.Response, err error) { fmt.Println("Request URL:", r.Request.URL, "failed with response:", r.Request, "\nError:", err) }) // 使用随机User-Agent发起请求 extensions.RandomUserAgent(c) c.Visit(site) return products }
main.go 原代码
package main import "github.com/Antonio-Costa00/Go-Price-Monitor/scraper" func main() { scraper.Scraper("https://sp.olx.com.br/?q=iphone%27") }
问题原因
- 异步执行逻辑问题:Colly的
c.Visit()是异步发起请求的,原代码在调用c.Visit()前就打印了products,此时回调函数还未执行,切片自然为空。 - 主函数提前返回:
c.Visit()不会阻塞主goroutine,主函数在所有回调完成前就返回了空切片。 - 并发数据竞争风险:Colly默认并发处理请求,多个回调goroutine同时修改切片会导致数据竞争,可能丢失数据。
修复方案
修改后的scraper.go代码
package scraper import ( "fmt" "strings" "sync" "github.com/gocolly/colly" "github.com/gocolly/colly/extensions" ) type Product struct { name string fullPrice string url string } func Scraper(site string) []Product { var mu sync.Mutex products := []Product{} c := colly.NewCollector() replacer := strings.NewReplacer("R$", "", ",", ".") c.OnHTML("div#column-main-content", func(e *colly.HTMLElement) { fullPrice := e.ChildText("span.m7nrfa-0.eJCbzj.sc-ifAKCX.ANnoQ") product := Product{ name: e.ChildText("h2"), fullPrice: replacer.Replace(fullPrice), url: e.ChildAttr("a.sc-1fcmfeb-2.iezWpY", "href"), } fmt.Println(product) // 加锁避免并发修改切片的数据竞争 mu.Lock() products = append(products, product) mu.Unlock() }) c.OnRequest(func(r *colly.Request) { fmt.Println("Visiting", r.URL) }) c.OnError(func(r *colly.Response, err error) { fmt.Println("Request URL:", r.Request.URL, "failed with response:", r.Request, "\nError:", err) }) extensions.RandomUserAgent(c) c.Visit(site) c.Wait() // 阻塞等待所有请求和回调执行完成 fmt.Println(products) return products }
关键修改点
- 调用
c.Wait():阻塞主goroutine,直到所有异步请求、回调处理完毕,确保切片填充完成后再返回。 - 添加
sync.Mutex互斥锁:避免多个并发回调goroutine同时修改切片导致的数据竞争问题。 - 调整
fmt.Println(products)位置:移到c.Wait()之后,此时切片已被正确填充,能输出有效数据。
内容的提问来源于stack exchange,提问作者Antonio Costa
相关产品推荐
相关产品推荐

