.NET Core(Linux)下保留纯净<p>标签并移除其余HTML标签的方案
Got it, let's work through this problem. Your current regex strips all HTML tags entirely, but you need to keep only clean <p></p> tags (no attributes) while removing every other tag like <div style="">, <span style="">, etc. Since you're on Linux with .NET Core and can't use full-framework parsers like AngleSharp, we can adjust your regex approach to fit this exact need.
Step 1: Fix the Regex Target
First, note that your current regex uses <.*?> which targets HTML-escaped tags (like <div>). If your input contains actual unescaped HTML tags (with < and >), you'll want to target <.*?> instead. Let's build the solution for unescaped HTML first (we'll cover the escaped case too).
Step 2: Two-Stage Regex Replacement
We'll use two regex passes to get the desired result:
- Convert any
<p>tag with attributes (like<p style="margin:0" class="foo">) to a clean<p>. - Strip all remaining HTML tags except the clean
<p>and</p>.
Here's the implementation:
using System.Text.RegularExpressions; public static string StripHtmlPreserveCleanPTags(string input) { // First pass: Replace any <p> with attributes to clean <p> // Matches <p followed by any non-> characters, then > string cleanPTags = Regex.Replace(input, @"<p\b[^>]*>", "<p>", RegexOptions.IgnoreCase); // Second pass: Strip all tags except clean <p> and </p> // Matches any tag that isn't exactly <p> or </p> string result = Regex.Replace(cleanPTags, @"<(?!\/?p\b)[^>]*>", string.Empty, RegexOptions.IgnoreCase); return result; }
How It Works
- First regex:
@"<p\b[^>]*>"<p\bmatches the start of a<p>tag (the word boundary\bensures we don't match tags like<pizza>)[^>]*matches any characters except>(all the attributes)- Replaces the entire attribute-rich
<p>tag with a clean<p>
- Second regex:
@"<(?!\/?p\b)[^>]*>"<matches the start of a tag(?!\/?p\b)is a negative lookahead: it skips tags that start withp(either<p>or</p>)[^>]*>matches the rest of the tag, which we replace with an empty string (strip it)
Handling HTML-Escaped Tags
If your input has escaped tags (like <p style="...">), adjust the regex to target the escaped entities:
public static string StripEscapedHtmlPreserveCleanPTags(string input) { // Clean escaped <p> tags string cleanPTags = Regex.Replace(input, @"<p\b[^&]*>", "<p>", RegexOptions.IgnoreCase); // Strip all other escaped tags string result = Regex.Replace(cleanPTags, @"<(?!\/?p\b)[^&]*>", string.Empty, RegexOptions.IgnoreCase); // Optional: Convert escaped <p> back to actual tags if needed // result = result.Replace("<p>", "<p>").Replace("</p>", "</p>"); return result; }
Test Example
Input:
<div style="color:red">Hello <p style="margin:0">World</p> <span>Test</span> <p class="bar">Another paragraph</p></div>
Output after processing:
Hello <p>World</p> Test <p>Another paragraph</p>
Limitations to Note
Regex isn't a perfect HTML parser, so this works best for:
- Standard content tags (no nested
<p>tags, which are invalid HTML anyway) - Tags that don't contain
>characters inside attribute values (e.g.,<p data-value=">">would break the regex)
If your HTML has edge cases like these, double-check if AngleSharp is actually compatible with .NET Core (it does support .NET Core now!)—it would be more reliable for complex HTML. But if that's a hard restriction, the regex approach will handle most common scenarios.
内容的提问来源于stack exchange,提问作者Patrick

