Solr未存储字段full_content缺失内容问题排查求助
full_content Values in Solr 7.6.0 First off, let's break down the possible reasons why your full_content fields are ending up empty in 200k out of 3 million documents—especially since your import process confirms the data is intact before sending to Solr. We'll cover analyzer behavior, import edge cases, memory constraints, and the impact of stored="false".
1. Analyzer Filtering Is Stripping All Content
Your text_general field type uses a chain of filters during indexing that could completely remove all tokens from your full_content value, leaving the field effectively empty in the index:
- LengthFilterFactory: Removes tokens shorter than 2 characters or longer than 100. If your content is made up entirely of single-character strings, this filter would wipe everything.
- StopFilterFactory: Removes common stopwords (from
stopwords.txt). If your content is nothing but stopwords (e.g., "a an the of"), this filter would eliminate all tokens. - Combine these, and even valid-looking content could get stripped down to nothing. Since
full_contentisstored="false", you can't directly check the raw value Solr received—you only know the field is empty because searches don't return the document.
2. Batch Import Logic Has Hidden Edge Cases
Even if your raw data looks good, your nightly import job might be accidentally clearing the full_content field when constructing SolrInputDocument objects:
- Maybe your code treats whitespace-only content as empty, or has a bug handling special characters (like non-UTF-8 encoding) that nulls out the field right before sending to Solr.
- Batch operations can sometimes have race conditions or partial failures that don't trigger obvious errors—especially if you're using bulk updates without strict error checking. Double-check the code that maps your source data to Solr fields, not just the raw source data itself.
3. Memory Constraints (Less Likely, But Possible)
Could memory issues cause this? It's not the most common culprit, but it's worth ruling out:
- If Solr's JVM heap is undersized, frequent garbage collection or memory pressure might interrupt the indexing process for some documents. In rare cases, this could lead to partial field processing (like tokenizers/filters failing to run, resulting in an empty field) without throwing an explicit OOM error.
- Check your Solr logs for GC warnings, and review your JVM settings (e.g.,
-Xmxvalue). If you see frequent full GC cycles, increasing the heap might help. That said, since you're not seeing any error logs, this is a lower-priority lead.
4. Does stored="false" Cause This Behavior?
No—stored="false" only means Solr doesn't save the raw field value for retrieval. It doesn't prevent the field from being indexed, as long as the analyzer produces tokens. The problem here is that the field is empty in the index, not just unretrievable.
That said, stored="false" makes troubleshooting harder because you can't verify the exact value Solr received.
Recommended Fixes & Next Steps
Temporarily Enable Storing for
full_content
Modify your schema to setstored="true"forfull_content, reload the schema, and run a small test import. Then query the documents to see if the stored value matches what you expected to send.- If the stored value is correct but searches still fail: Use Solr's Analysis Page to test your raw content against the
text_generalanalyzer. This will show you exactly which filters are stripping tokens. - If the stored value is empty: Your import job is sending empty values to Solr—dig deeper into your document construction logic.
- If the stored value is correct but searches still fail: Use Solr's Analysis Page to test your raw content against the
Validate Analyzer Behavior
Use the Analysis Page to input sample content that's supposed to be in the missing documents. Check if the index analyzer produces any tokens. If not, adjust your filters (e.g., expand the LengthFilter range, remove stopwords that are actually relevant to your content, or updatestopwords.txt).Check Import Job Error Handling
Ensure your nightly job is logging every step of document construction and Solr submission. Even silent failures (like invalid field values that Solr silently drops) might leave clues in detailed logs.Review Solr Memory Configuration
If you suspect memory issues, increase the JVM heap size (e.g., set-Xmx8gif you have available resources) and monitor GC logs to see if pressure decreases.
内容的提问来源于stack exchange,提问作者pierot

