Solr索引Flask搭建的HTML网站时出现404错误,求解Solr Cell相关配置问题
Hey there! Let's break down how to resolve this 404 error when using Solr Cell to index your Flask HTML site. You've already dug up the right docs—let's turn that info into actionable steps.
Why the 404 Happens
The data_driven_schema_configs default config set doesn't include Solr Cell (the module that handles extracting content from HTML/PDF/etc.) out of the box. That means the /update/extract endpoint doesn't exist yet, and the required jar files aren't loaded.
Step 1: Add Solr Cell Dependency Libraries
You need to add those <lib> tags to your solrconfig.xml (found in $SOLR_HOME/mycollection/conf/solrconfig.xml).
Place these lines inside the root <config> element—ideally near any existing <lib> entries (if there are none, put them right after the <luceneMatchVersion> line for consistency):
<lib dir="${solr.install.dir:../../..}/contrib/extraction/lib" regex=".*\.jar" /> <lib dir="${solr.install.dir:../../..}/dist/" regex="solr-cell-\d.*\.jar" />
These lines tell Solr to load all the jars needed for content extraction.
Step 2: Add the ExtractingRequestHandler
Your hunch is correct—you do need to add this request handler configuration.
Add this block inside the <config> element, usually alongside other <requestHandler> definitions (look for sections like <requestHandler name="/update" ...>):
<requestHandler name="/update/extract" startup="lazy" class="solr.extraction.ExtractingRequestHandler" > <lst name="defaults"> <str name="lowernames">true</str> <str name="fmap.content">_text_</str> </lst> </requestHandler>
This sets up the /update/extract endpoint with sensible defaults (like converting field names to lowercase and mapping extracted content to the _text_ field, which is indexed for search).
Step 3: Reload or Restart Solr
After saving your changes to solrconfig.xml, you need to make Solr pick up the new config. You can either:
- Restart the entire Solr service:
sudo su - solr -c "/opt/solr/bin/solr restart" - Or reload just your
mycollectioncore (faster):curl "http://localhost:8983/solr/admin/cores?action=RELOAD&core=mycollection"
Step 4: Test and Re-run Your Post Command
First, verify the endpoint exists by running:
curl http://localhost:8983/solr/mycollection/update/extract
You should get a 400 error (instead of 404) saying "Missing content stream"—that means the endpoint is now active!
Now re-run your post command to index your Flask site:
bin/post -c mycollection -filetypes html http://your-website-url
That should do it! You should no longer hit the 404, and Solr will start extracting and indexing your HTML content.
内容的提问来源于stack exchange,提问作者slash

