SolrCloud 6.2.1写入索引异常求助:分析错误致请求失败
Let's break down your recurring issue and walk through actionable steps to fix the Exception writing document id 999213794621579264 to the index; possible analysis error error. You’ve already tried several fixes (hard commits, static fields, JVM scaling, etc.), but the problem keeps coming back—here’s how to dig deeper:
1. Isolate the Problem Document First
The error points to a specific document ID, so start here to rule out content-related issues:
- Extract the document with ID
999213794621579264from your data source and try indexing it alone via Solr's API:curl -X POST -H "Content-Type: application/json" http://host:8983/solr/mycollection/update -d '[{"id":"999213794621579264", "field1":"value1", ...}]' - If this fails consistently, inspect every field in the document for red flags: non-UTF-8 characters, extremely long values, or data that doesn’t match the field’s type (e.g., a string in an integer field).
- Use Solr’s built-in analysis tool (under your collection’s "Analysis" tab in the admin UI) to test suspicious field values against their defined analyzers—this will reveal if the analysis chain is throwing an error.
2. Check for Corrupted Index Files
Restarting fixes the issue temporarily, which suggests index corruption might be happening during runtime:
- On each shard leader and replica, run Lucene’s
CheckIndextool to verify index integrity:java -jar solr-core-6.2.1.jar org.apache.lucene.index.CheckIndex /path/to/solr/data/mycollection/shardX/replicaN/index - If corruption is detected, repair it with the
-fixflag:java -jar solr-core-6.2.1.jar org.apache.lucene.index.CheckIndex /path/to/index -fix - Also, check your server’s disk health: look for IO errors, full disks, or slow IO in system logs (
/var/log/messagesordmesg). Poor disk performance can lead to partial writes and index corruption.
3. Refine Your Commit Strategy
Your current commit setup might be putting unnecessary stress on the system:
- Replace manual curl commits with Solr’s native auto-commit/soft-commit configuration in
solrconfig.xmlfor better control:<autoCommit> <maxTime>60000</maxTime> <!-- Auto hard commit every 1 minute --> <openSearcher>false</openSearcher> <!-- Avoid overhead of opening a new searcher --> </autoCommit> <autoSoftCommit> <maxTime>10000</maxTime> <!-- Soft commit every 10 seconds for near-real-time visibility --> </autoSoftCommit> - Reduce the frequency of hard commits (you’re doing one every 10 minutes now)—hard commits trigger expensive disk flushes. Try extending this to 30 minutes unless you need strict consistency.
- Limit concurrent bulk submit threads from your client tool—too many simultaneous writes can cause index lock conflicts or partial writes.
4. Tune JVM and GC Settings
Even with more heap, misconfigured GC can cause indexing timeouts:
- Use Solr-recommended JVM parameters for 6.2.1 (adjust heap size based on your server’s resources):
-Xms8g -Xmx8g -XX:+UseG1GC -XX:MaxGCPauseMillis=200 -XX:+ParallelRefProcEnabled -XX:+UnlockExperimentalVMOptions -XX:+DoEscapeAnalysis -XX:ParallelGCThreads=8 -XX:ConcGCThreads=2 -XX:InitiatingHeapOccupancyPercent=70 - Enable GC logging to identify long pauses or frequent full GCs that could block indexing:
-Xloggc:/var/log/solr/gc.log -XX:+PrintGCDetails -XX:+PrintGCDateStamps -XX:+PrintGCApplicationStoppedTime -XX:+UseGCLogFileRotation -XX:NumberOfGCLogFiles=5 -XX:GCLogFileSize=20m - Avoid setting heap size larger than 16GB unless you’re using a GC optimized for large heaps (like G1GC with proper tuning).
5. Verify SolrCloud Cluster Health
Shard replication issues can lead to inconsistent index states:
- Check the Cloud tab in Solr Admin to ensure all replicas are active and in sync with their leaders. Look for replicas marked as "down" or with high replication lag.
- If you see problematic replicas, delete and re-create them to refresh the index copy.
- Confirm your ZooKeeper cluster is stable—ZK failures can disrupt SolrCloud’s coordination and lead to index write errors.
6. Audit Field Definitions and Analyzers
Converting dynamic fields to static ones is good, but double-check for hidden mismatches:
- Ensure every field’s type matches the data you’re sending (e.g.,
intfields don’t get string values,datefields follow the correct format likeyyyy-MM-dd'T'HH:mm:ssZ). - If you’re using custom analyzers or third-party plugins, temporarily replace them with default analyzers (e.g.,
StandardTokenizerfor text fields) to see if the error goes away. This will help you pinpoint a buggy analysis component.
7. Consider Upgrading Solr
Solr 6.2.1 is over 7 years old and has known bugs related to indexing, concurrency, and analysis that have been fixed in newer versions:
- Upgrade to the latest 6.x release (6.6.6) first—it’s a drop-in replacement for 6.2.1 and includes critical bug fixes.
- If possible, move to a supported stable version like 8.x or 9.x for long-term stability and better performance.
内容的提问来源于stack exchange,提问作者user8955522

