You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取Episerver Content Area中所有块的XHTMLstring类型属性

提取Episerver内容区域中所有Block的XhtmlString属性文本

核心思路

通过反射遍历Block的所有属性,筛选出XhtmlString类型的属性,再借助Episerver内置的ToPlainText()方法提取纯文本;同时遍历Content Area中的所有内容项,批量处理每个Block。

具体实现代码

1. 通用提取单个Block的XhtmlString文本方法

using System;
using System.Collections.Generic;
using System.Reflection;
using EPiServer.Core;

public static class BlockTextExtractor
{
    // 提取单个IContent实例中所有XhtmlString属性的纯文本
    public static IEnumerable<string> ExtractXhtmlTextFromBlock(IContent block)
    {
        if (block == null)
            yield break;

        // 获取Block的所有公共实例属性
        var properties = block.GetType().GetProperties(BindingFlags.Public | BindingFlags.Instance);
        
        foreach (var prop in properties)
        {
            // 判断属性类型是否为XhtmlString
            if (prop.PropertyType == typeof(XhtmlString))
            {
                var xhtmlValue = prop.GetValue(block) as XhtmlString;
                if (xhtmlValue != null && !string.IsNullOrWhiteSpace(xhtmlValue.ToPlainText()))
                {
                    yield return xhtmlValue.ToPlainText().Trim();
                }
            }
            // 可选:递归处理嵌套的IContent类型属性
            else if (typeof(IContent).IsAssignableFrom(prop.PropertyType))
            {
                var nestedContent = prop.GetValue(block) as IContent;
                foreach (var text in ExtractXhtmlTextFromBlock(nestedContent))
                {
                    yield return text;
                }
            }
            // 可选:递归处理嵌套的ContentArea属性
            else if (prop.PropertyType == typeof(ContentArea))
            {
                var nestedArea = prop.GetValue(block) as ContentArea;
                if (nestedArea != null)
                {
                    foreach (var item in nestedArea.Items)
                    {
                        var nestedBlock = item.GetContent();
                        foreach (var text in ExtractXhtmlTextFromBlock(nestedBlock))
                        {
                            yield return text;
                        }
                    }
                }
            }
        }
    }

    // 遍历ContentArea中的所有Block,提取所有XhtmlString文本
    public static IEnumerable<string> ExtractAllXhtmlTextFromContentArea(ContentArea contentArea)
    {
        if (contentArea == null || contentArea.Items == null)
            yield break;

        foreach (var areaItem in contentArea.Items)
        {
            var block = areaItem.GetContent();
            if (block != null)
            {
                foreach (var text in ExtractXhtmlTextFromBlock(block))
                {
                    yield return text;
                }
            }
        }
    }
}

2. 使用示例

假设你的新闻页模型包含MainContentArea这个ContentArea属性,调用方式如下:

// currentPage为你的新闻页实例
var allText = BlockTextExtractor.ExtractAllXhtmlTextFromContentArea(currentPage.MainContentArea);
// 按需合并所有文本
var combinedText = string.Join(Environment.NewLine, allText);

注意事项

  • ToPlainText()是XhtmlString的内置方法,会自动剥离HTML标签,仅保留纯文本内容
  • 若你的项目中有自定义的XhtmlString子类,需要调整类型判断的条件
  • 递归处理逻辑为可选项,可根据Block是否存在嵌套内容结构决定是否启用

内容的提问来源于stack exchange,提问作者Ms1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 10:10:34