You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用GoQuery获取页面内容时,无法按换行(br)拆分元素求助

解决GoQuery无法按<br>标签拆分内容的问题

我明白你遇到的麻烦了——GoQuery默认的Text()方法会把所有文本内容揉成一段,完全无视<br>标签的换行作用,这确实挺闹心的。针对你的HTML结构,这里有两种靠谱的解决思路:

方法1:遍历子节点,手动拆分文本

这种方式直接操作DOM节点,能精准识别<br>标签并分割内容,适合需要精细控制的场景:

import (
    "fmt"
    "strings"
    "github.com/PuerkitoBio/goquery"
)

doc, err := goquery.NewDocumentFromReader(res.Body)
if err != nil {
    panic(err)
}

// 定位到目标p标签(第二个li里的内容容器)
doc.Find("ul li:nth-child(2) p").Each(func(i int, s *goquery.Selection) {
    var lines []string
    currentLine := ""

    // 遍历p标签下的所有子节点
    s.Contents().Each(func(j int, node *goquery.Selection) {
        if node.Is("br") {
            // 遇到br标签,把当前整理好的行存入切片并重置
            if trimmed := strings.TrimSpace(currentLine); trimmed != "" {
                lines = append(lines, trimmed)
                currentLine = ""
            }
        } else {
            // 文本节点,持续追加内容
            currentLine += node.Text()
        }
    })

    // 处理最后一行未被br收尾的内容
    if trimmed := strings.TrimSpace(currentLine); trimmed != "" {
        lines = append(lines, trimmed)
    }

    // 此时lines就是按br拆分好的每一行内容
    fmt.Println("拆分结果:", lines)
})

方法2:先转HTML替换br,再提取文本

如果不需要太细的节点控制,这种方法更简洁,通过替换HTML里的br标签来实现拆分:

import (
    "fmt"
    "strings"
    "github.com/PuerkitoBio/goquery"
)

doc, err := goquery.NewDocumentFromReader(res.Body)
if err != nil {
    panic(err)
}

doc.Find("ul li:nth-child(2) p").Each(func(i int, s *goquery.Selection) {
    // 获取p标签的原始HTML内容
    htmlContent, _ := s.Html()
    
    // 替换所有形式的br标签为换行符(兼容<br>、<br/>、<br />)
    htmlContent = strings.ReplaceAll(htmlContent, "<br>", "\n")
    htmlContent = strings.ReplaceAll(htmlContent, "<br/>", "\n")
    htmlContent = strings.ReplaceAll(htmlContent, "<br />", "\n")
    
    // 临时解析处理后的HTML,提取纯文本
    tempDoc, _ := goquery.NewDocumentFromReader(strings.NewReader(htmlContent))
    rawText := tempDoc.Text()
    
    // 按换行符拆分并清理空白行
    rawLines := strings.Split(rawText, "\n")
    cleanedLines := make([]string, 0, len(rawLines))
    for _, line := range rawLines {
        if trimmed := strings.TrimSpace(line); trimmed != "" {
            cleanedLines = append(cleanedLines, trimmed)
        }
    }

    fmt.Println("拆分结果:", cleanedLines)
})

小提示

  • 两种方法都先通过ul li:nth-child(2) p精准定位到目标内容容器,避免抓取到第一个li里的无关文本
  • 方法1的可控性更强,适合需要保留节点上下文的场景;方法2更轻量化,适合大多数普通需求

内容的提问来源于stack exchange,提问作者Belmin Bedak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 04:20:41