Chromedp获取UL子元素异常及批量爬取URL问题求助
问题1:Chromedp获取UL节点后无法遍历子LI元素
原因
chromedp通过chromedp.Nodes获取的节点默认不会递归加载子节点,因此即使ChildNodeCount显示有子元素,Children数组也会是空的。
解决方案
不需要依赖节点的Children属性,直接通过选择器或JS Eval获取子元素的文本:
方案1:直接批量获取所有LI下的文本
使用chromedp.TextAll直接定位到目标span元素,一次性获取所有文本:
var allSpecTexts []string err := chromedp.Run(ctx, chromedp.TextAll(`.product-item__content ul.product-small-specs li span`, &allSpecTexts), ) if err != nil { log.Fatal(err) } // 输出所有文本 for _, text := range allSpecTexts { fmt.Println(text) }
方案2:逐个UL处理子元素
如果需要对每个UL单独处理,先获取所有UL节点,再通过JS Eval获取每个UL下的LI文本:
var specs []*cdp.Node err := chromedp.Run(ctx, chromedp.Nodes(`.product-item__content ul.product-small-specs`, &specs, chromedp.AtLeast(0)), ) if err != nil { log.Fatal(err) } for _, ulNode := range specs { var liTexts []string // 执行JS获取当前UL下所有span的文本内容 err := chromedp.Run(ctx, chromedp.Eval(`(ul) => Array.from(ul.querySelectorAll('li span')).map(el => el.textContent.trim())`, ulNode, &liTexts), ) if err != nil { log.Printf("处理UL节点失败:%v", err) continue } fmt.Printf("当前UL的子元素文本:%v\n", liTexts) }
问题2:批量爬取URL时出现TypeError错误
错误原因
- 索引竞争问题:goroutine中使用循环变量
i会导致访问错误的节点(循环迭代速度快于goroutine启动,i已经被修改),即使传入了n,代码中仍使用links[i]而非links[n],可能访问到已失效的节点,调用AttributeValue时引发异常。 - 节点访问失效:页面导航后,之前获取的
links节点可能已被销毁,此时调用AttributeValue会导致不可预期的错误。 - 文本选择器问题:
chromedp.Text("h1")可能获取到非元素节点(比如文本节点),导致调用getClientRects方法失败。
解决方案
1. 修复goroutine索引问题
提前获取href值,或在goroutine中使用正确的索引变量:
maxGoroutines := 1 guard := make(chan struct{}, maxGoroutines) for i := range links { guard <- struct{}{} // 提前获取href,避免在goroutine中访问可能失效的节点 itemHref := links[i].AttributeValue("href") go func() { defer func() { <-guard }() // 确保释放guard retrieveDetails("https://www.bol.com" + itemHref) time.Sleep(5 * time.Second) }() }
2. 修正文本获取逻辑
避免获取非元素节点,改用JS Eval或添加节点可见性检查:
修改retrieveDetails中的文本获取步骤:
// 替换原chromedp.Text行 chromedp.Eval(`document.querySelector('h1')?.textContent.trim() || ''`, &header),
或者添加节点可见性约束:
chromedp.Text(`h1`, &header, chromedp.NodeVisible, chromedp.AtLeast(0)),
3. 优化资源使用(可选)
每次调用retrieveDetails创建新的ExecAllocator和Context会消耗大量资源,可复用ExecAllocator或限制并发数:
// 全局创建一次ExecAllocator opts := append(chromedp.DefaultExecAllocatorOptions[:], chromedp.Flag("headless", false), ) actx, acancel := chromedp.NewExecAllocator(context.Background(), opts...) defer acancel() // 在retrieveDetails中复用actx func retrieveDetails(actx context.Context, url string) { ctx, cancel := chromedp.NewContext(actx, chromedp.WithLogf(log.Printf)) defer cancel() ctx, cancel = context.WithTimeout(ctx, 6000*time.Second) defer cancel() // ... 其余代码不变 }
内容的提问来源于stack exchange,提问作者Angelo van Cleef
相关产品推荐
相关产品推荐

