You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Go访问HTML DOM并提取标签内容生成结构体?

需求可行性说明及实现示例

这个需求完全可行,核心流程分为三个环节:发起HTTP请求获取目标页面的HTML内容、解析HTML提取指定标签的文本、将提取结果赋值给自定义结构体。

以下用Go语言给出具体实现示例(其他语言如Python可通过对应库实现相同逻辑):

1. 定义自定义结构体

type Example struct {
    Foo string // 存储<h1>标签内容
    Bar string // 存储<p>标签内容
}

2. 完整实现代码

package main

import (
    "fmt"
    "io"
    "net/http"
    "strings"

    "golang.org/x/net/html"
)

type Example struct {
    Foo string
    Bar string
}

func main() {
    // 目标URL
    targetURL := "https://example.com/"
    
    // 发起HTTP请求获取HTML页面
    resp, err := http.Get(targetURL)
    if err != nil {
        fmt.Printf("请求页面失败: %v\n", err)
        return
    }
    defer resp.Body.Close()

    body, err := io.ReadAll(resp.Body)
    if err != nil {
        fmt.Printf("读取响应内容失败: %v\n", err)
        return
    }

    // 解析HTML文档
    doc, err := html.Parse(strings.NewReader(string(body)))
    if err != nil {
        fmt.Printf("解析HTML失败: %v\n", err)
        return
    }

    var result Example

    // 递归遍历HTML节点,提取目标标签内容
    var traverse func(*html.Node)
    traverse = func(n *html.Node) {
        if n.Type == html.ElementNode {
            switch n.Data {
            case "h1":
                // 提取<h1>的文本内容
                if n.FirstChild != nil {
                    result.Foo = n.FirstChild.Data
                }
            case "p":
                // 提取<p>下的所有文本内容并去除首尾空格
                var contentBuilder strings.Builder
                for child := n.FirstChild; child != nil; child = child.NextSibling {
                    if child.Type == html.TextNode {
                        contentBuilder.WriteString(child.Data)
                    }
                }
                result.Bar = strings.TrimSpace(contentBuilder.String())
            }
        }
        // 继续遍历子节点
        for child := n.FirstChild; child != nil; child = child.NextSibling {
            traverse(child)
        }
    }

    traverse(doc)

    // 输出最终结果
    fmt.Printf("提取完成:\nFoo(h1内容): %s\nBar(p内容): %s\n", result.Foo, result.Bar)
}

其他语言实现思路

如果使用Python,可借助requests库发起请求,BeautifulSoup库解析HTML,提取标签文本后赋值给自定义类(对应结构体),逻辑和上述示例完全一致。

内容的提问来源于stack exchange,提问作者irf98

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 18:09:24