制药行业长药品名多词片段反向匹配搜索实现方案咨询
It sounds like you need a lightweight way to match product names against unordered, fragment-based search queries (like reversed combinations or partial terms) without relying on full-text search engines or indexes. Here's a practical, step-by-step approach to implement this, including support for showing similar products:
1. Preprocess Product Data (One-Time Setup)
First, normalize your product catalog to make query comparisons straightforward:
Normalize Case: Convert all product names to lowercase to eliminate case sensitivity issues.
Split into Atomic Tokens: Break down each normalized name into meaningful, searchable chunks. Use regex to split on whitespace and non-alphanumeric characters, then separate numeric values from their units (e.g., "250MG/5ML" becomes ["250", "mg", "5", "ml"]).
Pseudocode example:
import re def tokenize_product_name(name): normalized = name.lower() # Extract alphanumeric chunks (ignores slashes, spaces, and special characters) return re.findall(r'[a-zA-Z]+|\d+', normalized)For
AMOXICILLIN SUSP 250MG/5ML 100ML, this returns:["amoxicillin", "susp", "250", "mg", "5", "ml", "100", "ml"]Group Similar Products: Create "drug groups" to handle the "同类产品" requirement. For example:
- Group all products containing variants of amoxicillin (like "amoxicillin", "amoxin", "amoxilin") into a single "amoxicillin-based" group. You can do this manually for small catalogs, or auto-group by checking if a product's tokens include any of your predefined related keywords.
2. Process Search Queries
For each incoming search query:
Normalize and Tokenize: Apply the same tokenization logic used for products. For example, the query
"100 250 sus amo"becomes["100", "250", "sus", "amo"].Filter Matching Products: Iterate through your product list and check if every query token exists as a substring in at least one of the product's normalized tokens. This handles unordered/reversed queries since we don't care about token order—only that all query fragments are present somewhere in the product's name.
Pseudocode for matching:
def is_product_match(product_tokens, query_tokens): for q_token in query_tokens: # Check if the query fragment exists in any product token match_found = any(q_token in p_token for p_token in product_tokens) if not match_found: return False return True
3. Include Similar Products
Once you have your matching products:
- Collect all unique drug groups from the matching products.
- Add every product in those groups to your results (even if they don't directly match the query). For example, if
AMOXICILLIN SUSP 250MG/5ML 100MLmatches, includeamoxilin 10mgsince it's in the same amoxicillin-based group.
4. Optional Optimizations
For larger product catalogs, linear scans can get slow. Here are lightweight tweaks to speed things up:
- Keyword Map: Precompute a dictionary where keys are common fragments (like "amo", "susp", "250") and values are lists of products containing those fragments. For a query, find the intersection of products across all query tokens to reduce the number of products you need to check.
- Early Termination: In the matching loop, stop checking a product as soon as one query token isn't found.
Example Walkthrough
Let’s test with query "amo sus 250 100":
- Query tokens:
["amo", "sus", "250", "100"] - Check
AMOXICILLIN SUSP 250MG/5ML 100ML:- "amo" is in "amoxicillin" → yes
- "sus" is in "susp" → yes
- "250" matches exactly → yes
- "100" matches exactly → yes → product is a match
- Check
AMOXIN SUSP 500MG/5ML 100ML:- "amo" is in "amoxin" → yes
- "sus" is in "susp" → yes
- "250" is not present → no → product doesn't match
- Add all products in the amoxicillin-based group (including
amoxilin 10mg) to the results.
This approach works without full-text indexes, handles unordered fragment queries, and includes similar products as required.
内容的提问来源于stack exchange,提问作者Prathap Gangireddy

