Go Colly爬取返回单个元素切片而非多元素切片问题求助
爬取网页时产品信息拼接问题的排查与解决
问题描述
我尝试爬取网页https://www.brasiltronic.com.br/pesquisa?pg=1&t=Fone%20de%20ouvido,使用以下scraper.go和main.go代码。运行后期望得到包含多个Product的切片,但实际返回仅含单个元素的切片,该元素的name、fullPrice字段为所有产品对应内容的拼接。
scraper.go代码
package scraper import ( "fmt" "strings" "github.com/gocolly/colly" "github.com/gocolly/colly/extensions" ) type Product struct { name string fullPrice string url string } func Scraper(url string) []Product { products := make([]Product, 0) c := colly.NewCollector() colly.AllowedDomains("www.brasiltronic.com.br") c.OnHTML("ul.row", func(e *colly.HTMLElement) { name := e.ChildText("div > div.information > h3.name.no-medium.no-tablet") fullPrice := e.ChildText("strong.sale-price > span:nth-child(1)") replacer := strings.NewReplacer("R$", "", ",", ".") fullPrice = replacer.Replace(fullPrice) url := e.ChildAttr("div > div.information > a", "href") products = append(products, Product{name: name, fullPrice: fullPrice, url: url}) }) c.OnError(func(r *colly.Response, err error) { fmt.Println("Request URL:", r.Request.URL, "failed with response:", r.Request, "\nError:", err) }) // Uses a random User-Agent in each request extensions.RandomUserAgent(c) c.Visit(url) return products }
main.go代码
package main import ( "fmt" "github.com/Antonio-Costa00/Go-Price-Monitor/scraper" ) func main() { url := "https://www.brasiltronic.com.br/pesquisa?pg=1&t=Fone%20e%20ouvido" products := scraper.Scraper(url) fmt.Println(products) }
运行输出
[{Fone de Ouvido Profissional AKG K92 com fio - Preto e DouradoMicrofone de lapela JBL com fone de ouvido CSLM 20Fone de Ouvido Samson SR350 Over-ear Estéreo PretoFone de Ouvido Sennheiser CX100 BrancoFone de Ouvido Audio -Technica ATH-M20xBT sem Fio com Bluetooth PretoFone de Ouvido Sennheiser HD100 com fio (Preto)Fone de Ouvido Sennheiser HD400S com fio (Preto)Fone de Ouvido Audio-Technica ATH-M40x Profissional para Monitoração com fio - PretoFone de Ouvido Audio-Technica ATH-M20x Profissional para Monitoração com fio - PretoFone de Ouvido Audi o-Technica ATH-M30x Profissional para Monitoração com fio - PretoFone de Ouvido Audio-Technica ATH-AVC400 extr a-auricualres SonicPro com fio - PretoFone de Ouvido Audio-Technica ATH-M50x Profissional para Monitoração com fio - PretoToca Discos Audio-Technica AT-LP60XHP-GM Automático Belt-Drive com Fone de Ouvido ATH-250AV ...Fon e de Ouvido Sem Fio Sennheiser RS2000 - PretoFone de Ouvido Profissional AKG K361-BT Dobrável - PretoKit Micro fone Samson C01U Pro PodCasting Pack SAC01UPROPKFone De Ouvido Profissional AKG K371-BT com Bluetooth - PretoH eadset Audio-Technica ATH-101USB Single-Ear com fio USB - PretoFone de Ouvido Audio-Technica ATH-R70x Profissi onal de referência abertos com fio - PretoKit Microfone Estudio Zoom ZUM-2 PMP com Headphone e tripé de mesaHe adset Audio-Technica ATH-102USB Dual-Ear com fio USB - PretoFones de ouvido de monitoramento Sennheiser IE 40 PRO Clear intra-auricularHeadset Gamer Audio-Technica ATH-G1WL Premium para Jogos Wireless - PretoMicrofone Sa mson Q9U Cardióide XLR/USB 359.10 94.50 159.30 159.30 699.30 269.10 519.30 879.30 419.40 619.20 373.50 1.219. 50 1.399.50 1.479.60 769.50 1.599.30 1.009.80 219.60 2.559.60 1.129.50 224.10 759.60 1.649.70 1.619.10 /fone-d e-ouvido-profissional-akg-k92-com-fio-preto-e-dourado-p1331225}]
问题原因
- AllowedDomains设置无效:
colly.AllowedDomains("www.brasiltronic.com.br")只是调用了函数但未将结果绑定到创建的Collector实例c上,这行代码不会起到限制域名的作用。 - 选择器匹配范围错误:
c.OnHTML的选择器是ul.row,这个节点是包含所有产品的整个列表容器。当调用e.ChildText时,Colly会把容器内所有匹配到的子元素文本拼接在一起,导致所有产品的名称、价格都合并成了一个字符串,最终只生成一个Product实例。
解决方案
1. 修正AllowedDomains配置
创建Collector时直接指定允许的域名,或者将无效的代码替换为对实例c的属性赋值:
// 方式一:创建Collector时指定 c := colly.NewCollector( colly.AllowedDomains("www.brasiltronic.com.br"), ) // 方式二:创建后赋值 c := colly.NewCollector() c.AllowedDomains = []string{"www.brasiltronic.com.br"}
2. 修改产品选择器
将OnHTML的选择器改为单个产品项的选择器。查看目标网页结构,ul.row下的每个li元素对应一个产品,因此将选择器改为ul.row li,这样回调函数会遍历每个产品项,逐个提取信息:
修正后的scraper.go代码
package scraper import ( "fmt" "strings" "github.com/gocolly/colly" "github.com/gocolly/colly/extensions" ) type Product struct { name string fullPrice string url string } func Scraper(url string) []Product { products := make([]Product, 0) // 创建Collector时指定允许的域名 c := colly.NewCollector( colly.AllowedDomains("www.brasiltronic.com.br"), ) // 选择单个产品项的节点 c.OnHTML("ul.row li", func(e *colly.HTMLElement) { name := e.ChildText("div > div.information > h3.name.no-medium.no-tablet") fullPrice := e.ChildText("strong.sale-price > span:nth-child(1)") replacer := strings.NewReplacer("R$", "", ",", ".") fullPrice = replacer.Replace(fullPrice) url := e.ChildAttr("div > div.information > a", "href") // 逐个添加产品到切片 products = append(products, Product{name: name, fullPrice: fullPrice, url: url}) }) c.OnError(func(r *colly.Response, err error) { fmt.Println("Request URL:", r.Request.URL, "failed with response:", r.Request, "\nError:", err) }) extensions.RandomUserAgent(c) c.Visit(url) return products }
说明
修改后,Collector会遍历每个产品项节点,分别提取每个产品的名称、价格和链接,最终生成包含多个Product实例的切片,符合预期结果。
内容的提问来源于stack exchange,提问作者Antonio Costa
相关产品推荐
相关产品推荐

