You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用HtmlAgilityPack抓取网页标题的技术问题求助

Fixing Your HtmlAgilityPack Web Scraping Issues in VB.NET

Let’s break down solutions for both of your scraping problems step by step.

Problem 1: Scraping Titles from nextinpact.com

Your initial code was almost there, but two key issues were holding it back:

  1. You used a full XPath (//h2...) which searches the entire document instead of limiting it to the target section.
  2. You were printing the section’s inner text instead of the individual title text for each result.

Here’s the corrected code:

Imports System
Imports System.Xml
Imports HtmlAgilityPack
Imports System.Net
Imports System.IO
Imports System.Collections.Generic

Public Class Program
    Public Shared Sub Main()
        ' Enable SSL Support
        ServicePointManager.SecurityProtocol = SecurityProtocolType.Tls12
        
        ' WebPage to scrape
        Dim link As String = "https://www.nextinpact.com"
        
        ' Download page into an HtmlDocument
        Dim doc As HtmlDocument = New HtmlWeb().Load(link)
        
        ' Select the target section
        Dim section As HtmlNode = doc.DocumentNode.SelectSingleNode("//section[@class='small_article_section']")
        
        If Not section Is Nothing Then
            ' Use relative XPath (".//") to search only within the selected section
            For Each titleLink As HtmlNode In section.SelectNodes(".//h2[@class='color_title']//a[@class='ui-link']")
                ' Extract and print the clean title text
                Console.WriteLine(titleLink.InnerText.Trim())
            Next
        Else
            Console.WriteLine("Target section not found on the page.")
        End If
    End Sub
End Class

Key Fixes:

  • Renamed div to section for clarity (since we’re targeting a <section> element).
  • Used .// in the XPath to restrict searches to nodes inside the selected section.
  • Updated the loop to print the inner text of each title link instead of the section’s text.

Problem 2: Scraping Product Titles from sideshow.com

The main issues here were syntax errors and incorrect text targeting:

  1. Using text() in your XPath only checks direct text nodes of the <h2>—if titles are nested in child elements (like <span>), this won’t capture them. Use . instead to check the full combined text of the <h2> node.
  2. Your XPath had a missing closing bracket, and you needed to combine both "Grace" and "Aspen" conditions with and.

Here’s the corrected code snippet:

Dim doc As HtmlDocument = New HtmlWeb().Load("https://www.sideshow.com/collectibles?manufacturer=sideshow+collectibles&type=premium+format%28tm%29+figure&brand=aspen")
Dim productList As HtmlNode = doc.DocumentNode.SelectSingleNode("//div[@class='c-ProductList row']")

If Not productList Is Nothing Then
    ' XPath to find h2 nodes containing both "Grace" and "Aspen" in their full text
    Dim targetTitles As HtmlNodeCollection = productList.SelectNodes(".//h2[contains(., 'Grace') and contains(., 'Aspen')]")
    
    If Not targetTitles Is Nothing Then
        For Each titleNode As HtmlNode In targetTitles
            Console.WriteLine(titleNode.InnerText.Trim())
        Next
    Else
        Console.WriteLine("No matching titles found.")
    End If
Else
    Console.WriteLine("Product list container not found on the page.")
End If

Key Fixes:

  • Replaced text() with . in contains() to capture all text within the <h2> (including child elements).
  • Combined both search terms using and to find titles that include both "Grace" and "Aspen".
  • Added null checks for the product list and title collection to avoid runtime errors.

Note on Dynamic Content:

If you still don’t see results, the product list might be loaded dynamically with JavaScript. HtmlAgilityPack only parses static HTML, so you’d need a tool like Selenium to render the page fully before scraping.

内容的提问来源于stack exchange,提问作者Manada

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:45:48