如何使用C#对包含多层转义&的文本进行彻底解码?
解决多层HTML转义的彻底解码问题
针对多层转义的HTML实体(比如&amp),单次WebUtility.HtmlDecode只能解码一层的问题,可以通过循环解码直到字符串不再变化的方式实现彻底解码。
实现代码
using System.Net; public static class HtmlDecodeHelper { public static string DecodeCompletely(string input) { if (string.IsNullOrEmpty(input)) return input; string current = input; string previous; do { previous = current; current = WebUtility.HtmlDecode(current); } while (current != previous); return current; } }
使用示例
调用这个方法处理你的测试字符串:
string MyString = "some text &amp"; string fullyDecoded = HtmlDecodeHelper.DecodeCompletely(MyString); // 输出结果:"some text &"
原理说明
每次调用HtmlDecode只会解码当前字符串中的所有顶层HTML实体,循环中不断重复解码操作,直到解码前后的字符串完全一致——这意味着已经没有可解码的实体存在,此时的结果就是彻底解码后的内容。这个方法同样适用于其他多层转义的HTML实体(比如&lt;会被解码为<)。
内容的提问来源于stack exchange,提问作者Udhayamani Vellaichamy
相关产品推荐
相关产品推荐

