WinForms C#网页数据提取求助:获取指定网页的文本与配料表内容
Hey there! Let's tackle this problem step by step—extracting content from that recipe page doesn't have to be a headache, even without IDs to target. Here's a solid approach that'll handle those multi-column ingredient lists and make your life easier:
Forget regex for HTML parsing—it's brittle and struggles with nested elements or variable layouts like multi-column lists. Instead, use HtmlAgilityPack, a .NET library that parses HTML into a navigable DOM tree, letting you target elements using class names, XPath, or CSS selectors.
Step 1: Install the Library
First, add HtmlAgilityPack to your project via NuGet:
- Using Package Manager Console:
Install-Package HtmlAgilityPack - Using .NET CLI:
dotnet add package HtmlAgilityPack
Step 2: Load the Web Page
Use the HtmlWeb class to fetch and parse the page. This handles basic HTTP requests and HTML parsing out of the box:
var web = new HtmlWeb(); var doc = web.Load("https://www.chefkoch.de/rezepte/drucken/512261146932016/Annas-Rouladen-mit-Seidenkloessen.html"); // Check if the document loaded successfully if (doc == null || doc.DocumentNode == null) { Console.WriteLine("Failed to load the page."); return; }
Step 3: Extract the Left-Side Text
Looking at the Chefkoch print page structure, the left-side recipe text is contained in a <div> with the class recipe-text. You can target it with an XPath selector:
// Get the left-side text container var leftTextNode = doc.DocumentNode.SelectSingleNode("//div[contains(@class, 'recipe-text')]"); if (leftTextNode != null) { // Extract clean text (or use InnerHtml if you need formatting) string leftSideText = leftTextNode.InnerText.Trim(); Console.WriteLine("Left-Side Recipe Text:\n" + leftSideText); } else { Console.WriteLine("Could not find the left-side text container."); }
Step 4: Extract Multi-Column Ingredients
The ingredient lists (even multi-column ones) use consistent class names—each ingredient item is in an <li> with the class ingredient. We can grab all these items regardless of which column they're in:
// Get all ingredient items var ingredientNodes = doc.DocumentNode.SelectNodes("//li[contains(@class, 'ingredient')]"); if (ingredientNodes != null) { Console.WriteLine("\nIngredients:"); foreach (var node in ingredientNodes) { // Extract the quantity and ingredient name (split if needed) string ingredientText = node.InnerText.Trim(); Console.WriteLine("- " + ingredientText); } } else { Console.WriteLine("Could not find any ingredients."); }
Why This Works for Multi-Column Lists
Chefkoch's multi-column ingredients are just multiple <ul> or <div> containers with the same ingredient item structure. By targeting the individual .ingredient list items directly, we bypass the column layout entirely—we don't care which column they're in, we just collect all of them.
Bonus: Handle Edge Cases
- If the page uses dynamic content (though the print page should be static), you might need a headless browser like PuppeteerSharp instead of HtmlWeb. But for Chefkoch's print view, HtmlWeb should work fine.
- If class names have variations (like
ingredient highlighted), thecontains(@class, 'ingredient')selector will still match, since it checks for the substring.
内容的提问来源于stack exchange,提问作者Michael

