You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为RawHTMLBlock添加search_index?FieldBlock无get_searchable_content属性能否搜索?

Answers to Your Wagtail Search Index Questions

1. Adding Search Index Support to RawHTMLBlock

By default, Wagtail's RawHTMLBlock doesn't feed content into the search index—since it stores raw HTML, the system can't easily extract usable text for search. The fix is to create a custom subclass that parses the HTML into plain text and exposes it to the search system.

Here's how to implement this:

First, install beautifulsoup4 (we'll use it to strip HTML tags):

pip install beautifulsoup4

Then define your custom block in blocks.py:

from bs4 import BeautifulSoup
from wagtail.core.blocks import RawHTMLBlock

class SearchableRawHTMLBlock(RawHTMLBlock):
    def get_searchable_content(self, value):
        # Convert HTML to plain text
        soup = BeautifulSoup(value, "html.parser")
        plain_text = soup.get_text(strip=True)
        # Return a list of strings (required for Wagtail's search API)
        return [plain_text] if plain_text else []

Replace the standard RawHTMLBlock with this custom version in your StreamField:

from wagtail.core.fields import StreamField

class MyPage(Page):
    body = StreamField([
        ("html_section", SearchableRawHTMLBlock()),
        # ... other blocks in your field
    ])

Now any content in your HTML blocks will be converted to plain text and included in the page's search index.

2. Searching FieldBlock Instances (Without get_searchable_content)

You’re correct that basic FieldBlock subclasses (like TextFieldBlock or CharBlock) don’t have a built-in get_searchable_content method—but you absolutely can make them searchable. Here are two simple approaches:

Option 1: Create a Searchable FieldBlock Subclass

Mirror the RawHTMLBlock solution by extending the FieldBlock and adding the missing method:

from wagtail.core.blocks import TextFieldBlock

class SearchableTextFieldBlock(TextFieldBlock):
    def get_searchable_content(self, value):
        # Return the field's value as a list (required format)
        return [value] if value else []

Use this custom block in your StreamField, and its content will automatically be picked up by the search index.

Option 2: Manually Include Content in the Page’s Searchable Content

If you don’t want to create a custom block, override your Page model’s get_searchable_content method to pull content directly from FieldBlock entries:

class MyPage(Page):
    body = StreamField([
        ("plain_text", TextFieldBlock()),
        # ... other blocks
    ])

    def get_searchable_content(self):
        # Start with the default searchable content (page title, etc.)
        content = super().get_searchable_content()
        
        # Loop through all blocks in the StreamField
        for block in self.body:
            # Target the specific FieldBlock type you want to index
            if block.block_type == "plain_text":
                # Add the block's value to the search content list
                if block.value:
                    content.append(block.value)
        
        return content

This gives you precise control over which blocks contribute to the search index.


内容的提问来源于stack exchange,提问作者Nick Mamashin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 03:52:42