Elasticsearch字段拼接匹配实现咨询:根据输入返回对应member_id
First off, let's confirm: your requirement is totally achievable in ES, and there are two main approaches depending on your performance and data update needs. Let's break them down step by step.
Option 1: Dynamic calculation during query (good for small datasets or frequent updates)
This approach computes the concatedData on the fly when you run your query, so you don't need to modify your existing document structure. Here's how to do it:
Elasticsearch DSL Query Example
You'll use a script query to concatenate the fields and compare against your keyInput:
GET /your_index/_search { "query": { "bool": { "filter": { "script": { "source": "doc['e_id'].value.toString() + doc['c_id'].value.toString() + Math.ceil(doc['salary'].value).toString() == params.keyInput", "params": { "keyInput": "YOUR_INPUT_STRING_HERE" } } } } }, "_source": ["member_id"] // Only return the member_id field we need }
Java Implementation (using RestHighLevelClient)
Here's how you'd translate this into Java code:
import org.elasticsearch.action.search.SearchRequest; import org.elasticsearch.action.search.SearchResponse; import org.elasticsearch.client.RequestOptions; import org.elasticsearch.client.RestHighLevelClient; import org.elasticsearch.index.query.BoolQueryBuilder; import org.elasticsearch.index.query.QueryBuilders; import org.elasticsearch.index.query.ScriptQueryBuilder; import org.elasticsearch.script.Script; import org.elasticsearch.script.ScriptType; import java.io.IOException; import java.util.HashMap; import java.util.Map; public class EsSearchExample { public void findMemberId(RestHighLevelClient client, String indexName, String keyInput) throws IOException { // Build the script with parameters Map<String, Object> params = new HashMap<>(); params.put("keyInput", keyInput); Script script = new Script(ScriptType.INLINE, "painless", "doc['e_id'].value.toString() + doc['c_id'].value.toString() + Math.ceil(doc['salary'].value).toString() == params.keyInput", params); // Build the filter query BoolQueryBuilder boolQuery = QueryBuilders.boolQuery() .filter(new ScriptQueryBuilder(script)); // Create search request and specify to only return member_id SearchRequest searchRequest = new SearchRequest(indexName); searchRequest.source().query(boolQuery).fetchSource(new String[]{"member_id"}, null); // Execute the request and process results SearchResponse response = client.search(searchRequest, RequestOptions.DEFAULT); response.getHits().forEach(hit -> { String memberId = hit.getSourceAsMap().get("member_id").toString(); System.out.println("Matched member_id: " + memberId); }); } }
Option 2: Precompute the concatedData field (better for large datasets/frequent queries)
If you're dealing with a lot of data or run this query often, precomputing the concatenated field during indexing will be much faster (script queries can be slow on big datasets). Here's how to set this up:
Step 1: Update your index mapping (add the new field)
First, add a concatedData field of type keyword (since we need exact matches):
PUT /your_index/_mapping { "properties": { "concatedData": { "type": "keyword" } } }
Step 2: Populate the field (two ways)
A. Compute during document indexing (in your Java app)
When you're preparing the document to index, calculate concatedData upfront:
// Assume you have a Member object with your fields Member member = ...; String concatedData = member.getE_id().toString() + member.getC_id().toString() + String.valueOf(Math.ceil(member.getSalary())); // Add the computed field to your document map Map<String, Object> doc = new HashMap<>(); doc.put("member_id", member.getMember_id()); doc.put("e_id", member.getE_id()); doc.put("c_id", member.getC_id()); doc.put("salary", member.getSalary()); doc.put("concatedData", concatedData); // Index the document as you normally would
B. Use an Ingest Pipeline (let Elasticsearch handle the calculation)
If you don't want to modify your Java code, create an ingest pipeline that automatically computes concatedData when documents are indexed:
PUT /_ingest/pipeline/concate_data_pipeline { "processors": [ { "script": { "source": "ctx.concatedData = ctx.e_id.toString() + ctx.c_id.toString() + Math.ceil(ctx.salary).toString()" } } ] }
Then specify this pipeline when indexing documents in Java:
IndexRequest indexRequest = new IndexRequest("your_index") .id(documentId) .source(doc) .setPipeline("concate_data_pipeline"); client.index(indexRequest, RequestOptions.DEFAULT);
Step 3: Query the precomputed field
Now your query becomes simple and fast (no scripts needed):
GET /your_index/_search { "query": { "term": { "concatedData": "YOUR_INPUT_STRING_HERE" } }, "_source": ["member_id"] }
The Java code for this query uses a basic termQuery, which is far more efficient than a script-based query.
Key Notes
- Ensure data type consistency: since
e_idandc_idare numbers, you must convert them to strings before concatenation to avoid unexpected behavior. - For
salary,Math.ceil()works correctly on float values in Painless (Elasticsearch's built-in scripting language). - If you choose the precomputed approach, don't forget to reindex existing documents to populate the
concatedDatafield for them.
内容的提问来源于stack exchange,提问作者MishraJi

