VB.NET从网页提取商品信息并在RichTextBox生成表格
VB.NET实现网页商品数据提取并生成表格到RichTextBox
嘿,我来帮你搞定这个需求!从网页源码里提取商品信息并生成指定格式的表格到RichTextBox,用专门的HTML解析库会比硬写正则靠谱太多——毕竟正则处理HTML很容易踩各种格式不规范的坑。下面一步步来:
第一步:安装HTML解析库HtmlAgilityPack
首先得给项目装个HtmlAgilityPack,这是.NET生态里处理HTML最常用的工具之一,能轻松定位和提取节点内容。安装方式很简单:
- 打开Visual Studio的NuGet包管理器(右键项目→管理NuGet包)
- 搜索
HtmlAgilityPack,点击安装即可
第二步:核心代码实现
先在代码文件顶部引用命名空间:
Imports HtmlAgilityPack
接下来写提取逻辑,比如放在一个按钮的点击事件里(你也可以根据自己的需求调整触发时机):
Private Sub btnExtractAndShow_Click(sender As Object, e As EventArgs) Handles btnExtractAndShow.Click ' 1. 准备网页源码:这里用变量模拟,实际可以从网络下载或本地读取 ' 如果是从网络获取,后面会给你补充示例代码 Dim htmlSource As String = "替换成你的网页源码内容" ' 2. 初始化HTML解析对象 Dim htmlDoc As New HtmlDocument() htmlDoc.LoadHtml(htmlSource) ' 3. 定位所有商品节点:选择所有class为aditem的article标签 Dim adItemNodes As HtmlNodeCollection = htmlDoc.DocumentNode.SelectNodes("//article[@class='aditem']") ' 4. 拼接结果文本 Dim resultContent As New StringBuilder() ' 添加表头(如果需要的话) resultContent.AppendLine("名称 | 价格 | 商品链接") resultContent.AppendLine("-------------------------------") ' 用横线模拟分隔线,适配RichTextBox If adItemNodes IsNot Nothing Then For Each itemNode As HtmlNode In adItemNodes ' 提取商品名称:定位aditem-main下的a标签文本 Dim nameNode = itemNode.SelectSingleNode(".//div[@class='aditem-main']//h2//a") Dim itemName = If(nameNode IsNot Nothing, nameNode.InnerText.Trim(), "无商品名称") ' 提取商品价格:定位aditem-details下的strong标签文本 Dim priceNode = itemNode.SelectSingleNode(".//div[@class='aditem-details']//strong") Dim itemPrice = If(priceNode IsNot Nothing, priceNode.InnerText.Trim(), "无价格") ' 提取商品链接:获取a标签的href属性值 Dim linkNode = itemNode.SelectSingleNode(".//div[@class='aditem-main']//h2//a") Dim itemLink = If(linkNode IsNot Nothing AndAlso linkNode.HasAttributes, linkNode.GetAttributeValue("href", "无链接"), "无链接") ' 拼接成指定格式的一行 resultContent.AppendLine($"{itemName} | {itemPrice} | {itemLink}") Next Else resultContent.AppendLine("未找到任何商品数据") End If ' 5. 将结果显示到RichTextBox1 RichTextBox1.Text = resultContent.ToString() ' 可选:设置等宽字体让表格更整齐 RichTextBox1.Font = New Font("Consolas", 9) End Sub
额外实用提示
- 从网络获取网页源码:如果你的源码是从在线网页获取的,可以用
WebClient或者HttpClient,示例代码如下:
' 用WebClient下载(简单场景) Dim wc As New WebClient() wc.Encoding = System.Text.Encoding.UTF8 ' 匹配网页编码,避免乱码 Dim htmlSource As String = wc.DownloadString("你的目标网页URL")
- 处理动态加载内容:如果网页是滚动加载或者用JS渲染的静态源码里没有数据,那可能需要用Selenium这类工具模拟浏览器加载,但如果是纯静态源码,上面的代码就足够了
- 空值判断:代码里的
If...IsNot Nothing是为了避免网页结构不规范导致的空引用错误,很重要哦
内容的提问来源于stack exchange,提问作者SeoMain Seo
相关产品推荐
相关产品推荐

