基于词语级编辑距离的Elasticsearch模糊语句搜索可行性问询
Absolutely, you can implement this kind of query in Elasticsearch. The core idea is to use script scoring with Elasticsearch's Painless scripting language to calculate word-level edit distances between your query and document content, then rank results by how closely they match.
Here's a step-by-step breakdown:
1. Prepare your index with tokenized fields
First, you need to store the tokenized (split into individual words) version of your expression field so the script can easily access individual words. You can do this with an ingest pipeline that splits the text into a word array during indexing.
Create the ingest pipeline
PUT _ingest/pipeline/tokenize_expression { "processors": [ { "analyze": { "field": "expression", "target_field": "expression_tokens", "analyzer": "standard" } } ] }
This pipeline uses the standard analyzer to split your expression text into words and stores them in a new expression_tokens array field.
Create your index (with the pipeline attached)
PUT /expressions_index { "settings": { "default_pipeline": "tokenize_expression" }, "mappings": { "properties": { "expression": { "type": "text", "analyzer": "standard" }, "expression_tokens": { "type": "keyword", "fielddata": true } } } }
Now when you index documents, the expression_tokens field will automatically be populated with the split words from expression.
2. Run the script-scored query
Use a script_score query to calculate the word-level edit distance between your query and each document. We'll use Painless's built-in edit_distance function to measure character-level distance between individual words, then sum those distances to get an overall score (lower total distance = higher relevance score).
For your example query "tell me something on elasticsearch", first get its tokenized version (via the standard analyzer: ["tell", "me", "something", "on", "elasticsearch"]), then use this query:
POST /expressions_index/_search { "query": { "script_score": { // First filter down to likely matches to improve performance "query": { "match": { "expression": "tell me something" } }, "script": { "source": """ def queryTokens = params.queryTokens; def docTokens = doc['expression_tokens']; if (docTokens.isEmpty()) { return 0; } int totalDistance = 0; // For each word in the query, find the closest word in the document for (def qToken : queryTokens) { int minDist = Integer.MAX_VALUE; for (def dToken : docTokens) { int dist = edit_distance(qToken, dToken); if (dist < minDist) { minDist = dist; } } totalDistance += minDist; } // Invert the total distance so closer matches get higher scores return 1.0 / (totalDistance + 1); """, "params": { "queryTokens": ["tell", "me", "something", "on", "elasticsearch"] } } } }, "sort": [ { "_score": { "order": "desc" } } ] }
How this works for your example
- For the document
{"expression": "tell me something about elasticsearch"}:
Most words are exact matches (distance 0), only"on"vs"about"has an edit distance of 3. Total distance is 3, so score is1/(3+1) = 0.25. - For the document
{"expression": "tell me something about kibana"}:
The words"tell","me","something"are exact matches,"on"vs"about"has distance 3, and"elasticsearch"vs"kibana"has a higher distance. Even so, it will still rank higher than unrelated documents, and you'll see it in your results.
Key notes
- Performance: Script scoring can be slow on large indexes, so always use a filter query (like the
matchin the example) to narrow down candidate documents first. - Customize the logic: If you want a strict word-sequence edit distance (counting word additions/deletions/replacements instead of character-level distances), you can implement a dynamic programming algorithm in Painless to calculate that—though it's more complex.
- Thresholds: You can add a
min_scoreparameter to filter out results that are too far from your query.
内容的提问来源于stack exchange,提问作者Swapnil Masurekar

