如何用XPath选取表头含Posición的表格所有行(C#实现)
How to Scrape a Wikipedia Table with a Specific Header Text Using C# and XPath
Hey there! As someone who's tackled Wikipedia table scraping in C# before, let me break down exactly how to get this done for you—since you're new to XPath, I'll keep things clear and actionable.
Step 1: Use the Right Library
First, you'll need a reliable HTML parsing library. HtmlAgilityPack is the go-to choice for C#; it handles messy HTML (like Wikipedia's) perfectly and supports XPath queries seamlessly.
Install it via NuGet:
- Using Package Manager Console:
Install-Package HtmlAgilityPack - Using .NET CLI:
dotnet add package HtmlAgilityPack
Step 2: Complete Working Code Example
Here's a full, commented script that fetches the target Wikipedia page, locates your desired table, and extracts all its rows:
using System; using System.Net.Http; using HtmlAgilityPack; class WikipediaTableScraper { static async System.Threading.Tasks.Task Main(string[] args) { // Replace this with your target Wikipedia article URL string wikiUrl = "https://es.wikipedia.org/wiki/Your_Target_Article"; // Fetch the page content using (HttpClient client = new HttpClient()) { try { string htmlContent = await client.GetStringAsync(wikiUrl); // Load HTML into a parseable document HtmlDocument doc = new HtmlDocument(); doc.LoadHtml(htmlContent); // XPath query to find tables with a <th> containing "Posición" // Adding @class='wikitable' narrows results to standard content tables (recommended) string xpathQuery = "//table[contains(@class, 'wikitable') and .//th[contains(normalize-space(text()), 'Posición')]]"; HtmlNode targetTable = doc.DocumentNode.SelectSingleNode(xpathQuery); if (targetTable == null) { Console.WriteLine("Could not find the table with 'Posición' header."); return; } // Extract all rows from the target table HtmlNodeCollection rows = targetTable.SelectNodes(".//tr"); if (rows == null) { Console.WriteLine("No rows found in the target table."); return; } // Loop through rows and print cell content foreach (HtmlNode row in rows) { HtmlNodeCollection cells = row.SelectNodes(".//th | .//td"); if (cells != null) { foreach (HtmlNode cell in cells) { // Clean up messy whitespace and print cell text string cellText = cell.InnerText.Trim().Replace("\n", " "); Console.Write($"{cellText}\t"); } Console.WriteLine(); } } } catch (Exception ex) { Console.WriteLine($"Error occurred: {ex.Message}"); } } } }
Key XPath Breakdown
Let's demystify the XPath query so you can tweak it later:
//table: Finds all<table>elements on the pagecontains(@class, 'wikitable'): Filters for Wikipedia's standard content tables (avoids matching sidebar/navigation tables).//th[contains(normalize-space(text()), 'Posición')]: Looks inside the table for a<th>element where normalized text (removes extra spaces/newlines) contains "Posición"normalize-space()is critical here—Wikipedia often has line breaks or extra spaces in header HTML, so this ensures you don't miss matches.
Pro Tips for Reliability
- Exact Match Option: If you know the header text is exactly "Posición" (no extra spaces), swap
contains(normalize-space(text()), 'Posición')withnormalize-space(text())='Posición'for stricter matching. - Handle HTML Changes: Wikipedia occasionally updates its table structure. If your query stops working, right-click the table in your browser and inspect the HTML to adjust the XPath.
- Consider the Wikipedia API: For long-term projects, the Wikipedia API is more stable than scraping HTML—you can request table data directly in structured formats like JSON. But XPath is perfect for quick one-off tasks!
- Expand Error Handling: The example includes basic error handling, but you can add checks for missing cells, network timeouts, or invalid URLs to make it more robust.
内容的提问来源于stack exchange,提问作者IvanHid
相关产品推荐
相关产品推荐

