You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C#中使用Regex提取HTML标签内文本的技术求助

Solution for Extracting Text from
Tags in HTML

Hey there! I see you've been spending hours trying to get regex to pull text from those <h1> tags, and you've managed to grab "meow" but need the rest. Let's break this down and fix it, plus talk about a better approach than regex for HTML handling.

First: Fixing Your Regex to Get All Matches

Your current regex is on the right track, but if you're only getting the first "meow", it's probably because you're grabbing a single Match instead of all Matches. Also, let's tweak the regex a bit to handle whitespace better and make sure it captures all content inside <h1> tags (even with nested elements like that <span>).

Here's the adjusted code to get all <h1> contents:

string html = @"<div class=""ewok-rater-header-section""> <ul class=""header""> <li><h1>meow</h1></li> <li><h1>meow2</h1></li> <li><h1>Time = <span class=""work-weight"">9.0 minutes</span></h1></li> </ul> </div>";

var regexpost = new System.Text.RegularExpressions.Regex(@"<h1\s*[^>]*>(.*?)</h1>", System.Text.RegularExpressions.RegexOptions.Singleline);
var matches = regexpost.Matches(html);

foreach (Match match in matches)
{
    // This will give you the full content inside each <h1>, including nested tags
    string h1Content = match.Groups[1].Value;
    Console.WriteLine(h1Content);
}

Let's explain the regex tweaks:

  • <h1\s*[^>]*>: Matches the opening <h1> tag, accounting for any whitespace (\s*) and any attributes ([^>]* is safer than .*? for cleanly stopping at the tag's closing >)
  • RegexOptions.Singleline: Makes the . match newlines, which is a safe addition in case your HTML has line breaks between tags

Running this will output:

meow
meow2
Time = <span class="work-weight">9.0 minutes</span>

But Wait: Regex Isn't Great for HTML

If you want the pure text (without the <span> tags in the third <h1>), regex gets messy fast. HTML isn't a regular language, so regex can break easily if the HTML structure changes (like extra line breaks, nested tags, or unexpected attributes).

A much better approach is to use an HTML parser library like HtmlAgilityPack (it's the go-to tool for parsing HTML in C#). Here's how to do it:

First, install HtmlAgilityPack via NuGet (right-click your project > Manage NuGet Packages > search for HtmlAgilityPack and install).

Then use this code:

string html = @"<div class=""ewok-rater-header-section""> <ul class=""header""> <li><h1>meow</h1></li> <li><h1>meow2</h1></li> <li><h1>Time = <span class=""work-weight"">9.0 minutes</span></h1></li> </ul> </div>";

HtmlDocument doc = new HtmlDocument();
doc.LoadHtml(html);

// Get all <h1> tags using XPath
var h1Tags = doc.DocumentNode.SelectNodes("//h1");

if (h1Tags != null)
{
    foreach (var h1 in h1Tags)
    {
        // Get the inner text (automatically strips all HTML tags)
        string pureText = h1.InnerText.Trim();
        Console.WriteLine(pureText);
    }
}

This will output the clean, tag-free text you probably want:

meow
meow2
Time = 9.0 minutes

This approach is way more robust—if the HTML structure changes slightly (like adding a class to an <h1> or rearranging whitespace), it will still work perfectly.

内容的提问来源于stack exchange,提问作者Andreaa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:42:25