C#中使用Regex提取HTML标签内文本的技术求助
Hey there! I see you've been spending hours trying to get regex to pull text from those <h1> tags, and you've managed to grab "meow" but need the rest. Let's break this down and fix it, plus talk about a better approach than regex for HTML handling.
First: Fixing Your Regex to Get All Matches
Your current regex is on the right track, but if you're only getting the first "meow", it's probably because you're grabbing a single Match instead of all Matches. Also, let's tweak the regex a bit to handle whitespace better and make sure it captures all content inside <h1> tags (even with nested elements like that <span>).
Here's the adjusted code to get all <h1> contents:
string html = @"<div class=""ewok-rater-header-section""> <ul class=""header""> <li><h1>meow</h1></li> <li><h1>meow2</h1></li> <li><h1>Time = <span class=""work-weight"">9.0 minutes</span></h1></li> </ul> </div>"; var regexpost = new System.Text.RegularExpressions.Regex(@"<h1\s*[^>]*>(.*?)</h1>", System.Text.RegularExpressions.RegexOptions.Singleline); var matches = regexpost.Matches(html); foreach (Match match in matches) { // This will give you the full content inside each <h1>, including nested tags string h1Content = match.Groups[1].Value; Console.WriteLine(h1Content); }
Let's explain the regex tweaks:
<h1\s*[^>]*>: Matches the opening<h1>tag, accounting for any whitespace (\s*) and any attributes ([^>]*is safer than.*?for cleanly stopping at the tag's closing>)RegexOptions.Singleline: Makes the.match newlines, which is a safe addition in case your HTML has line breaks between tags
Running this will output:
meow meow2 Time = <span class="work-weight">9.0 minutes</span>
But Wait: Regex Isn't Great for HTML
If you want the pure text (without the <span> tags in the third <h1>), regex gets messy fast. HTML isn't a regular language, so regex can break easily if the HTML structure changes (like extra line breaks, nested tags, or unexpected attributes).
A much better approach is to use an HTML parser library like HtmlAgilityPack (it's the go-to tool for parsing HTML in C#). Here's how to do it:
First, install HtmlAgilityPack via NuGet (right-click your project > Manage NuGet Packages > search for HtmlAgilityPack and install).
Then use this code:
string html = @"<div class=""ewok-rater-header-section""> <ul class=""header""> <li><h1>meow</h1></li> <li><h1>meow2</h1></li> <li><h1>Time = <span class=""work-weight"">9.0 minutes</span></h1></li> </ul> </div>"; HtmlDocument doc = new HtmlDocument(); doc.LoadHtml(html); // Get all <h1> tags using XPath var h1Tags = doc.DocumentNode.SelectNodes("//h1"); if (h1Tags != null) { foreach (var h1 in h1Tags) { // Get the inner text (automatically strips all HTML tags) string pureText = h1.InnerText.Trim(); Console.WriteLine(pureText); } }
This will output the clean, tag-free text you probably want:
meow meow2 Time = 9.0 minutes
This approach is way more robust—if the HTML structure changes slightly (like adding a class to an <h1> or rearranging whitespace), it will still work perfectly.
内容的提问来源于stack exchange,提问作者Andreaa

