You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

.NET Core(Linux)下保留纯净<p>标签并移除其余HTML标签的方案

Solution for Preserving Clean

Tags While Stripping Other HTML in .NET Core (Linux)

Got it, let's work through this problem. Your current regex strips all HTML tags entirely, but you need to keep only clean <p></p> tags (no attributes) while removing every other tag like <div style="">, <span style="">, etc. Since you're on Linux with .NET Core and can't use full-framework parsers like AngleSharp, we can adjust your regex approach to fit this exact need.

Step 1: Fix the Regex Target

First, note that your current regex uses &lt;.*?&gt; which targets HTML-escaped tags (like &lt;div&gt;). If your input contains actual unescaped HTML tags (with < and >), you'll want to target <.*?> instead. Let's build the solution for unescaped HTML first (we'll cover the escaped case too).

Step 2: Two-Stage Regex Replacement

We'll use two regex passes to get the desired result:

  1. Convert any <p> tag with attributes (like <p style="margin:0" class="foo">) to a clean <p>.
  2. Strip all remaining HTML tags except the clean <p> and </p>.

Here's the implementation:

using System.Text.RegularExpressions;

public static string StripHtmlPreserveCleanPTags(string input)
{
    // First pass: Replace any <p> with attributes to clean <p>
    // Matches <p followed by any non-> characters, then >
    string cleanPTags = Regex.Replace(input, @"<p\b[^>]*>", "<p>", RegexOptions.IgnoreCase);
    
    // Second pass: Strip all tags except clean <p> and </p>
    // Matches any tag that isn't exactly <p> or </p>
    string result = Regex.Replace(cleanPTags, @"<(?!\/?p\b)[^>]*>", string.Empty, RegexOptions.IgnoreCase);
    
    return result;
}

How It Works

  • First regex: @"<p\b[^>]*>"
    • <p\b matches the start of a <p> tag (the word boundary \b ensures we don't match tags like <pizza>)
    • [^>]* matches any characters except > (all the attributes)
    • Replaces the entire attribute-rich <p> tag with a clean <p>
  • Second regex: @"<(?!\/?p\b)[^>]*>"
    • < matches the start of a tag
    • (?!\/?p\b) is a negative lookahead: it skips tags that start with p (either <p> or </p>)
    • [^>]*> matches the rest of the tag, which we replace with an empty string (strip it)

Handling HTML-Escaped Tags

If your input has escaped tags (like &lt;p style="..."&gt;), adjust the regex to target the escaped entities:

public static string StripEscapedHtmlPreserveCleanPTags(string input)
{
    // Clean escaped <p> tags
    string cleanPTags = Regex.Replace(input, @"&lt;p\b[^&]*&gt;", "&lt;p&gt;", RegexOptions.IgnoreCase);
    // Strip all other escaped tags
    string result = Regex.Replace(cleanPTags, @"&lt;(?!\/?p\b)[^&]*&gt;", string.Empty, RegexOptions.IgnoreCase);
    // Optional: Convert escaped <p> back to actual tags if needed
    // result = result.Replace("&lt;p&gt;", "<p>").Replace("&lt;/p&gt;", "</p>");
    return result;
}

Test Example

Input:

<div style="color:red">Hello <p style="margin:0">World</p> <span>Test</span> <p class="bar">Another paragraph</p></div>

Output after processing:

Hello <p>World</p> Test <p>Another paragraph</p>

Limitations to Note

Regex isn't a perfect HTML parser, so this works best for:

  • Standard content tags (no nested <p> tags, which are invalid HTML anyway)
  • Tags that don't contain > characters inside attribute values (e.g., <p data-value=">"> would break the regex)
    If your HTML has edge cases like these, double-check if AngleSharp is actually compatible with .NET Core (it does support .NET Core now!)—it would be more reliable for complex HTML. But if that's a hard restriction, the regex approach will handle most common scenarios.

内容的提问来源于stack exchange,提问作者Patrick

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 14:27:32