使用HtmlAgilityPack抓取网页标题的技术问题求助
Fixing Your HtmlAgilityPack Web Scraping Issues in VB.NET
Let’s break down solutions for both of your scraping problems step by step.
Problem 1: Scraping Titles from nextinpact.com
Your initial code was almost there, but two key issues were holding it back:
- You used a full XPath (
//h2...) which searches the entire document instead of limiting it to the target section. - You were printing the section’s inner text instead of the individual title text for each result.
Here’s the corrected code:
Imports System Imports System.Xml Imports HtmlAgilityPack Imports System.Net Imports System.IO Imports System.Collections.Generic Public Class Program Public Shared Sub Main() ' Enable SSL Support ServicePointManager.SecurityProtocol = SecurityProtocolType.Tls12 ' WebPage to scrape Dim link As String = "https://www.nextinpact.com" ' Download page into an HtmlDocument Dim doc As HtmlDocument = New HtmlWeb().Load(link) ' Select the target section Dim section As HtmlNode = doc.DocumentNode.SelectSingleNode("//section[@class='small_article_section']") If Not section Is Nothing Then ' Use relative XPath (".//") to search only within the selected section For Each titleLink As HtmlNode In section.SelectNodes(".//h2[@class='color_title']//a[@class='ui-link']") ' Extract and print the clean title text Console.WriteLine(titleLink.InnerText.Trim()) Next Else Console.WriteLine("Target section not found on the page.") End If End Sub End Class
Key Fixes:
- Renamed
divtosectionfor clarity (since we’re targeting a<section>element). - Used
.//in the XPath to restrict searches to nodes inside the selected section. - Updated the loop to print the inner text of each title link instead of the section’s text.
Problem 2: Scraping Product Titles from sideshow.com
The main issues here were syntax errors and incorrect text targeting:
- Using
text()in your XPath only checks direct text nodes of the<h2>—if titles are nested in child elements (like<span>), this won’t capture them. Use.instead to check the full combined text of the<h2>node. - Your XPath had a missing closing bracket, and you needed to combine both "Grace" and "Aspen" conditions with
and.
Here’s the corrected code snippet:
Dim doc As HtmlDocument = New HtmlWeb().Load("https://www.sideshow.com/collectibles?manufacturer=sideshow+collectibles&type=premium+format%28tm%29+figure&brand=aspen") Dim productList As HtmlNode = doc.DocumentNode.SelectSingleNode("//div[@class='c-ProductList row']") If Not productList Is Nothing Then ' XPath to find h2 nodes containing both "Grace" and "Aspen" in their full text Dim targetTitles As HtmlNodeCollection = productList.SelectNodes(".//h2[contains(., 'Grace') and contains(., 'Aspen')]") If Not targetTitles Is Nothing Then For Each titleNode As HtmlNode In targetTitles Console.WriteLine(titleNode.InnerText.Trim()) Next Else Console.WriteLine("No matching titles found.") End If Else Console.WriteLine("Product list container not found on the page.") End If
Key Fixes:
- Replaced
text()with.incontains()to capture all text within the<h2>(including child elements). - Combined both search terms using
andto find titles that include both "Grace" and "Aspen". - Added null checks for the product list and title collection to avoid runtime errors.
Note on Dynamic Content:
If you still don’t see results, the product list might be loaded dynamically with JavaScript. HtmlAgilityPack only parses static HTML, so you’d need a tool like Selenium to render the page fully before scraping.
内容的提问来源于stack exchange,提问作者Manada
相关产品推荐
相关产品推荐

