You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

WinForms C#网页数据提取求助:获取指定网页的文本与配料表内容

Hey there! Let's tackle this problem step by step—extracting content from that recipe page doesn't have to be a headache, even without IDs to target. Here's a solid approach that'll handle those multi-column ingredient lists and make your life easier:

Forget regex for HTML parsing—it's brittle and struggles with nested elements or variable layouts like multi-column lists. Instead, use HtmlAgilityPack, a .NET library that parses HTML into a navigable DOM tree, letting you target elements using class names, XPath, or CSS selectors.

Step 1: Install the Library

First, add HtmlAgilityPack to your project via NuGet:

  • Using Package Manager Console: Install-Package HtmlAgilityPack
  • Using .NET CLI: dotnet add package HtmlAgilityPack

Step 2: Load the Web Page

Use the HtmlWeb class to fetch and parse the page. This handles basic HTTP requests and HTML parsing out of the box:

var web = new HtmlWeb();
var doc = web.Load("https://www.chefkoch.de/rezepte/drucken/512261146932016/Annas-Rouladen-mit-Seidenkloessen.html");

// Check if the document loaded successfully
if (doc == null || doc.DocumentNode == null)
{
    Console.WriteLine("Failed to load the page.");
    return;
}

Step 3: Extract the Left-Side Text

Looking at the Chefkoch print page structure, the left-side recipe text is contained in a <div> with the class recipe-text. You can target it with an XPath selector:

// Get the left-side text container
var leftTextNode = doc.DocumentNode.SelectSingleNode("//div[contains(@class, 'recipe-text')]");

if (leftTextNode != null)
{
    // Extract clean text (or use InnerHtml if you need formatting)
    string leftSideText = leftTextNode.InnerText.Trim();
    Console.WriteLine("Left-Side Recipe Text:\n" + leftSideText);
}
else
{
    Console.WriteLine("Could not find the left-side text container.");
}

Step 4: Extract Multi-Column Ingredients

The ingredient lists (even multi-column ones) use consistent class names—each ingredient item is in an <li> with the class ingredient. We can grab all these items regardless of which column they're in:

// Get all ingredient items
var ingredientNodes = doc.DocumentNode.SelectNodes("//li[contains(@class, 'ingredient')]");

if (ingredientNodes != null)
{
    Console.WriteLine("\nIngredients:");
    foreach (var node in ingredientNodes)
    {
        // Extract the quantity and ingredient name (split if needed)
        string ingredientText = node.InnerText.Trim();
        Console.WriteLine("- " + ingredientText);
    }
}
else
{
    Console.WriteLine("Could not find any ingredients.");
}

Why This Works for Multi-Column Lists

Chefkoch's multi-column ingredients are just multiple <ul> or <div> containers with the same ingredient item structure. By targeting the individual .ingredient list items directly, we bypass the column layout entirely—we don't care which column they're in, we just collect all of them.

Bonus: Handle Edge Cases

  • If the page uses dynamic content (though the print page should be static), you might need a headless browser like PuppeteerSharp instead of HtmlWeb. But for Chefkoch's print view, HtmlWeb should work fine.
  • If class names have variations (like ingredient highlighted), the contains(@class, 'ingredient') selector will still match, since it checks for the substring.

内容的提问来源于stack exchange,提问作者Michael

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 19:37:27