You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用XPath选取表头含Posición的表格所有行(C#实现)

How to Scrape a Wikipedia Table with a Specific Header Text Using C# and XPath

Hey there! As someone who's tackled Wikipedia table scraping in C# before, let me break down exactly how to get this done for you—since you're new to XPath, I'll keep things clear and actionable.

Step 1: Use the Right Library

First, you'll need a reliable HTML parsing library. HtmlAgilityPack is the go-to choice for C#; it handles messy HTML (like Wikipedia's) perfectly and supports XPath queries seamlessly.

Install it via NuGet:

  • Using Package Manager Console:
    Install-Package HtmlAgilityPack
    
  • Using .NET CLI:
    dotnet add package HtmlAgilityPack
    

Step 2: Complete Working Code Example

Here's a full, commented script that fetches the target Wikipedia page, locates your desired table, and extracts all its rows:

using System;
using System.Net.Http;
using HtmlAgilityPack;

class WikipediaTableScraper
{
    static async System.Threading.Tasks.Task Main(string[] args)
    {
        // Replace this with your target Wikipedia article URL
        string wikiUrl = "https://es.wikipedia.org/wiki/Your_Target_Article";

        // Fetch the page content
        using (HttpClient client = new HttpClient())
        {
            try
            {
                string htmlContent = await client.GetStringAsync(wikiUrl);

                // Load HTML into a parseable document
                HtmlDocument doc = new HtmlDocument();
                doc.LoadHtml(htmlContent);

                // XPath query to find tables with a <th> containing "Posición"
                // Adding @class='wikitable' narrows results to standard content tables (recommended)
                string xpathQuery = "//table[contains(@class, 'wikitable') and .//th[contains(normalize-space(text()), 'Posición')]]";
                
                HtmlNode targetTable = doc.DocumentNode.SelectSingleNode(xpathQuery);

                if (targetTable == null)
                {
                    Console.WriteLine("Could not find the table with 'Posición' header.");
                    return;
                }

                // Extract all rows from the target table
                HtmlNodeCollection rows = targetTable.SelectNodes(".//tr");

                if (rows == null)
                {
                    Console.WriteLine("No rows found in the target table.");
                    return;
                }

                // Loop through rows and print cell content
                foreach (HtmlNode row in rows)
                {
                    HtmlNodeCollection cells = row.SelectNodes(".//th | .//td");
                    if (cells != null)
                    {
                        foreach (HtmlNode cell in cells)
                        {
                            // Clean up messy whitespace and print cell text
                            string cellText = cell.InnerText.Trim().Replace("\n", " ");
                            Console.Write($"{cellText}\t");
                        }
                        Console.WriteLine();
                    }
                }
            }
            catch (Exception ex)
            {
                Console.WriteLine($"Error occurred: {ex.Message}");
            }
        }
    }
}

Key XPath Breakdown

Let's demystify the XPath query so you can tweak it later:

  • //table: Finds all <table> elements on the page
  • contains(@class, 'wikitable'): Filters for Wikipedia's standard content tables (avoids matching sidebar/navigation tables)
  • .//th[contains(normalize-space(text()), 'Posición')]: Looks inside the table for a <th> element where normalized text (removes extra spaces/newlines) contains "Posición"
    • normalize-space() is critical here—Wikipedia often has line breaks or extra spaces in header HTML, so this ensures you don't miss matches.

Pro Tips for Reliability

  • Exact Match Option: If you know the header text is exactly "Posición" (no extra spaces), swap contains(normalize-space(text()), 'Posición') with normalize-space(text())='Posición' for stricter matching.
  • Handle HTML Changes: Wikipedia occasionally updates its table structure. If your query stops working, right-click the table in your browser and inspect the HTML to adjust the XPath.
  • Consider the Wikipedia API: For long-term projects, the Wikipedia API is more stable than scraping HTML—you can request table data directly in structured formats like JSON. But XPath is perfect for quick one-off tasks!
  • Expand Error Handling: The example includes basic error handling, but you can add checks for missing cells, network timeouts, or invalid URLs to make it more robust.

内容的提问来源于stack exchange,提问作者IvanHid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:05:22