You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Chromedp获取UL子元素异常及批量爬取URL问题求助

问题1:Chromedp获取UL节点后无法遍历子LI元素

原因

chromedp通过chromedp.Nodes获取的节点默认不会递归加载子节点,因此即使ChildNodeCount显示有子元素,Children数组也会是空的。

解决方案

不需要依赖节点的Children属性,直接通过选择器或JS Eval获取子元素的文本:

方案1:直接批量获取所有LI下的文本

使用chromedp.TextAll直接定位到目标span元素,一次性获取所有文本:

var allSpecTexts []string
err := chromedp.Run(ctx,
    chromedp.TextAll(`.product-item__content ul.product-small-specs li span`, &allSpecTexts),
)
if err != nil {
    log.Fatal(err)
}
// 输出所有文本
for _, text := range allSpecTexts {
    fmt.Println(text)
}

方案2:逐个UL处理子元素

如果需要对每个UL单独处理,先获取所有UL节点,再通过JS Eval获取每个UL下的LI文本:

var specs []*cdp.Node
err := chromedp.Run(ctx,
    chromedp.Nodes(`.product-item__content ul.product-small-specs`, &specs, chromedp.AtLeast(0)),
)
if err != nil {
    log.Fatal(err)
}

for _, ulNode := range specs {
    var liTexts []string
    // 执行JS获取当前UL下所有span的文本内容
    err := chromedp.Run(ctx,
        chromedp.Eval(`(ul) => Array.from(ul.querySelectorAll('li span')).map(el => el.textContent.trim())`, ulNode, &liTexts),
    )
    if err != nil {
        log.Printf("处理UL节点失败:%v", err)
        continue
    }
    fmt.Printf("当前UL的子元素文本:%v\n", liTexts)
}

问题2:批量爬取URL时出现TypeError错误

错误原因

  1. 索引竞争问题:goroutine中使用循环变量i会导致访问错误的节点(循环迭代速度快于goroutine启动,i已经被修改),即使传入了n,代码中仍使用links[i]而非links[n],可能访问到已失效的节点,调用AttributeValue时引发异常。
  2. 节点访问失效:页面导航后,之前获取的links节点可能已被销毁,此时调用AttributeValue会导致不可预期的错误。
  3. 文本选择器问题:chromedp.Text("h1")可能获取到非元素节点(比如文本节点),导致调用getClientRects方法失败。

解决方案

1. 修复goroutine索引问题

提前获取href值,或在goroutine中使用正确的索引变量:

maxGoroutines := 1
guard := make(chan struct{}, maxGoroutines)

for i := range links {
    guard <- struct{}{}
    // 提前获取href,避免在goroutine中访问可能失效的节点
    itemHref := links[i].AttributeValue("href")
    go func() {
        defer func() { <-guard }() // 确保释放guard
        retrieveDetails("https://www.bol.com" + itemHref)
        time.Sleep(5 * time.Second)
    }()
}

2. 修正文本获取逻辑

避免获取非元素节点,改用JS Eval或添加节点可见性检查:
修改retrieveDetails中的文本获取步骤:

// 替换原chromedp.Text行
chromedp.Eval(`document.querySelector('h1')?.textContent.trim() || ''`, &header),

或者添加节点可见性约束:

chromedp.Text(`h1`, &header, chromedp.NodeVisible, chromedp.AtLeast(0)),

3. 优化资源使用(可选)

每次调用retrieveDetails创建新的ExecAllocator和Context会消耗大量资源,可复用ExecAllocator或限制并发数:

// 全局创建一次ExecAllocator
opts := append(chromedp.DefaultExecAllocatorOptions[:],
    chromedp.Flag("headless", false),
)
actx, acancel := chromedp.NewExecAllocator(context.Background(), opts...)
defer acancel()

// 在retrieveDetails中复用actx
func retrieveDetails(actx context.Context, url string) {
    ctx, cancel := chromedp.NewContext(actx, chromedp.WithLogf(log.Printf))
    defer cancel()

    ctx, cancel = context.WithTimeout(ctx, 6000*time.Second)
    defer cancel()

    // ... 其余代码不变
}

内容的提问来源于stack exchange,提问作者Angelo van Cleef

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 19:30:49