Solr v5及以上版本中,判断字段存在的*与[* TO *]语法差异咨询
Great question! Let’s dive into the key differences between my_field:* and my_field:[* TO *] in Solr 5 and later—these two syntaxes might look interchangeable at first, but they behave differently depending on field types, performance, and edge cases.
Key Differences Between my_field:* and my_field:[* TO *]
1. Compatibility with Field/Data Types
my_field:*(Wildcard Query)
This wildcard match works best with string fields—it’ll catch any non-empty string value stored in the field. However, it falls flat with numeric (int, long), boolean, or date fields. Solr’s wildcard queries target terms in the inverted index, and numeric/boolean fields (especially those using PointFields, the standard in Solr 5+) don’t store values in a way that wildcard queries can reliably match. You might get no results or unexpected behavior if you use this on non-string fields.my_field:[* TO *](Range Query)
This range query is built to work with all field types that support range operations—that includes numerics, booleans, dates, and strings. For numeric fields, it matches any valid non-null value; for booleans, it catches bothtrueandfalse; for strings, it matches any non-empty string (since it covers the full range of lex ordered terms). It’s the consistent, reliable choice across different field types.
2. Handling of Empty/Null Values
Both syntaxes will not match documents where the field is missing entirely or has a null value. There’s a subtle edge case with empty strings too: if your string field allows empty strings and they’re indexed, neither query will match them (Solr typically doesn’t index empty string terms, so they fall outside both the wildcard and range match criteria). So both behave the same here—only non-empty, existing field values are picked up.
3. Performance
my_field:*
Wildcard queries ending with*are notoriously slow in Solr. They force the index to scan every term in the field’s inverted index to find matches, which is brutal for high-cardinality fields (like user IDs or product names). For large indexes, this can cause significant latency or even timeouts.my_field:[* TO *]
Range queries are optimized in Solr, especially with PointFields (the default for numeric/datetime fields in Solr 5+). They leverage the index’s sorted structure to quickly locate matching documents without scanning every term. Even for string fields, range queries are more efficient than wildcards because they use lex ordering to narrow down matches faster.
4. Multi-Value Fields
For multi-value fields, both queries will match a document if at least one value in the field is non-empty. The same type compatibility rules apply though: my_field:* might fail to match numeric multi-value fields, while [* TO *] will work as expected.
Recommendation
If you need to check for the presence of any non-null value in a field—regardless of its type—use my_field:[* TO *]. It’s reliable across all field types, performs better, and avoids the pitfalls of wildcard queries. Reserve my_field:* only for specific cases where you’re targeting string fields and can accept the performance hit.
内容的提问来源于stack exchange,提问作者Tomer Arazy

