You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Go语言colly框架爬取相同类名元素及背景图URL的实现方法

Go Colly 爬虫字段提取修正方案

原有代码问题

原有代码存在两个核心问题:

  1. 遍历逻辑错误:遍历对象为.list span即所有list下的全部span节点,且循环内使用外层element而非当前循环的elem取值,导致每次都拿到所有list下第二个span的文本拼接
  2. 缺少字段拆分和属性解析逻辑:所有结果存入同一数组,也没有处理style属性的正则解析提取背景图URL

修正后代码

package main

import (
    "fmt"
    "regexp"
    "github.com/gocolly/colly/v2"
)

// 预编译正则匹配background-image的url内容
var bgURLRegex = regexp.MustCompile(`background-image:\s*url\((.*?)\)`)

func main() {
    c := colly.NewCollector()

    // 初始化三个存储切片
    var (
        countrybg []string
        continet  []string
        country   []string
    )

    c.OnHTML(".cc", func(element *colly.HTMLElement) {
        // 遍历每个.list节点,保证三个字段一一对应
        element.ForEach(".list", func(_ int, listItem *colly.HTMLElement) {
            // 提取背景图URL
            bgStyle, exist := listItem.DOM.Find(".countrybg").Attr("style")
            if exist {
                matchRes := bgURLRegex.FindStringSubmatch(bgStyle)
                if len(matchRes) == 2 {
                    countrybg = append(countrybg, matchRes[1])
                }
            }
            // 提取大洲名称
            continet = append(continet, listItem.ChildText(".continet"))
            // 提取国家名称
            country = append(country, listItem.ChildText(".country"))
        })

        // 输出符合要求的结果
        fmt.Printf("countrybg = %q\n", countrybg)
        fmt.Printf("continet = %q\n", continet)
        fmt.Printf("country = %q\n", country)
    })

    c.Visit("你的目标页面URL")
}

提示:如果目标页面的外层容器属性确实拼写为clas="cc"(少了一个s),请将OnHTML的选择器修改为div[clas="cc"]即可匹配。

关键修改说明

  • 调整遍历层级为逐个遍历.list节点,保证每个节点下的三个字段是一一对应的,不会出现顺序错乱
  • 直接通过类名定位对应span,不需要依赖节点顺序,兼容性更强
  • 用预编译的正则表达式提取style属性中的背景图地址,自动过滤url()包裹格式
  • 三个字段分别存入独立切片,直接输出即可得到你需要的格式

内容的提问来源于stack exchange,提问作者Dinesh s

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 02:06:00