如何为RawHTMLBlock添加search_index?FieldBlock无get_searchable_content属性能否搜索?
1. Adding Search Index Support to RawHTMLBlock
By default, Wagtail's RawHTMLBlock doesn't feed content into the search index—since it stores raw HTML, the system can't easily extract usable text for search. The fix is to create a custom subclass that parses the HTML into plain text and exposes it to the search system.
Here's how to implement this:
First, install beautifulsoup4 (we'll use it to strip HTML tags):
pip install beautifulsoup4
Then define your custom block in blocks.py:
from bs4 import BeautifulSoup from wagtail.core.blocks import RawHTMLBlock class SearchableRawHTMLBlock(RawHTMLBlock): def get_searchable_content(self, value): # Convert HTML to plain text soup = BeautifulSoup(value, "html.parser") plain_text = soup.get_text(strip=True) # Return a list of strings (required for Wagtail's search API) return [plain_text] if plain_text else []
Replace the standard RawHTMLBlock with this custom version in your StreamField:
from wagtail.core.fields import StreamField class MyPage(Page): body = StreamField([ ("html_section", SearchableRawHTMLBlock()), # ... other blocks in your field ])
Now any content in your HTML blocks will be converted to plain text and included in the page's search index.
2. Searching FieldBlock Instances (Without get_searchable_content)
You’re correct that basic FieldBlock subclasses (like TextFieldBlock or CharBlock) don’t have a built-in get_searchable_content method—but you absolutely can make them searchable. Here are two simple approaches:
Option 1: Create a Searchable FieldBlock Subclass
Mirror the RawHTMLBlock solution by extending the FieldBlock and adding the missing method:
from wagtail.core.blocks import TextFieldBlock class SearchableTextFieldBlock(TextFieldBlock): def get_searchable_content(self, value): # Return the field's value as a list (required format) return [value] if value else []
Use this custom block in your StreamField, and its content will automatically be picked up by the search index.
Option 2: Manually Include Content in the Page’s Searchable Content
If you don’t want to create a custom block, override your Page model’s get_searchable_content method to pull content directly from FieldBlock entries:
class MyPage(Page): body = StreamField([ ("plain_text", TextFieldBlock()), # ... other blocks ]) def get_searchable_content(self): # Start with the default searchable content (page title, etc.) content = super().get_searchable_content() # Loop through all blocks in the StreamField for block in self.body: # Target the specific FieldBlock type you want to index if block.block_type == "plain_text": # Add the block's value to the search content list if block.value: content.append(block.value) return content
This gives you precise control over which blocks contribute to the search index.
内容的提问来源于stack exchange,提问作者Nick Mamashin

