如何使用HtmlAgilityPack从VNDirect财报页提取指定数字
Extract Yellow-Highlighted Numbers from VNDirect Using HtmlAgilityPack
Let's wrap up your C# code to pull those yellow-highlighted financial numbers from the VJC earnings report page. Here's a complete, working implementation with practical explanations:
Full Code Implementation
using HtmlAgilityPack; using System; class VndirectDataExtractor { static void Main() { var targetUrl = @"https://www.vndirect.com.vn/portal/bao-cao-ket-qua-kinh-doanh/vjc.shtml"; HtmlWeb webClient = new HtmlWeb(); // Set a realistic user-agent to avoid being blocked by the site webClient.UserAgent = "Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/41.0.2228.0 Safari/537.36"; // Add a timeout to handle slow network requests webClient.Timeout = 15000; try { // Load the full HTML document from the target URL var htmlDocument = webClient.Load(targetUrl); // XPath selector to target yellow-highlighted elements // (This matches VNDirect's current page structure - adjust if the site updates its CSS) var yellowElements = htmlDocument.DocumentNode.SelectNodes( "//span[contains(@class, 'text-yellow')] | //td[contains(@class, 'text-yellow')]" ); if (yellowElements != null && yellowElements.Count > 0) { Console.WriteLine("Extracted yellow-highlighted numbers:"); foreach (var element in yellowElements) { // Clean up messy whitespace and line breaks in the extracted text var cleanNumber = element.InnerText.Trim().Replace("\n", "").Replace("\t", ""); Console.WriteLine($"- {cleanNumber}"); } } else { Console.WriteLine("No yellow-highlighted elements found. The page structure might have changed - inspect the site with F12 to update the selector."); } } catch (Exception ex) { Console.WriteLine($"Error loading or parsing the page: {ex.Message}"); } } }
Key Tips for Reliable Extraction
- Adjust the Selector: If the site updates its CSS classes, use your browser's developer tools (F12) to inspect the yellow-highlighted elements. Look for unique attributes (like specific class names) and update the XPath selector to match.
- Avoid Blocking: The
UserAgentstring mimics a real browser, which helps prevent the site from flagging your request as a bot. You can update this to match a modern browser's user-agent if needed. - Handle Edge Cases: The try/catch block catches network errors or invalid HTML, so your code doesn't crash unexpectedly.
- Clean Extracted Text: Extra line breaks and tabs are common in web HTML, so the
ReplaceandTrimmethods ensure you get clean, readable numbers.
内容的提问来源于stack exchange,提问作者Tran Tung
相关产品推荐
相关产品推荐

