Watson Discovery Service结构化数据查询:模糊搜索及实体识别咨询
Great question! Let's break this down into the two core issues you're facing—getting fuzzy search to work reliably, and training the service to recognize similar entities like name variants.
Fuzzy Search Support in Watson Discovery
First, let's clarify how fuzzy search works in the service, and why your initial attempts didn't give the results you wanted:
Why Your Queries Failed
firstname:jon: This is an exact term match, so it only returns documents where thefirstnamefield contains the exact token "jon". Since your target document has "john", it won't show up.firstname::!jon: The::!operator disables exact matching, but it doesn't enable character-level fuzzy matching. Instead, it loosens the match based on the field's analyzer settings (like lowercase normalization, tokenization) but doesn't account for typos or minor character differences. That's why it returned unrelated records.
Correct Fuzzy Search Methods
Use these approaches to get targeted fuzzy matches for your example:
- Wildcard searches: Use
*(matches multiple characters) or?(matches a single character) to match partial terms:firstname:jo*nwill match "john", "jon", "jordan" (adjust the wildcard placement to narrow results if needed)firstname:jon?will match "john" (the?replaces the missing "h")
- Fuzzy matching operator: Use the tilde
~to allow for a specific number of character edits (insertions, deletions, substitutions):firstname:jon~(default edit distance of 2) will match "john" (1 edit difference) and other close variantsfirstname:jon~1limits matches to terms with exactly 1 character difference, which is more precise for your use case
Training Watson Discovery to Recognize Entity Variants
If you want consistent recognition of similar entities (like "john" and "jon" as the same first name), here are the most effective methods:
1. Custom Dictionaries
Create a custom dictionary to group variant terms as equivalent. This tells Discovery to treat all terms in a group as interchangeable during searches.
- How to set it up:
- Go to your Discovery collection in the console
- Navigate to Customization > Dictionaries
- Click "Create dictionary" and name it something like "FirstNameVariants"
- Add term groups: e.g., one group containing "john", "jon", "johnny", "jonathan"
- Save the dictionary and apply it to your collection
Now, searching for any term in the group will return documents matching any of the variants.
2. Query Expansion Training
Train Discovery to associate specific queries with relevant documents, so it learns to expand searches to include related terms over time.
- How to set it up:
- Go to Training > Query Expansion in your collection
- Add a query like "jon"
- Select documents that are relevant (e.g., those with "john" in the
firstnamefield) and mark them as "Relevant" - Repeat with a few more examples (e.g., query "john" and mark documents with "jon" as relevant)
The system will gradually learn to connect these variants and return more accurate results.
3. Custom Entity Recognition (Document Understanding)
If you're using the Document Understanding feature, you can train a custom entity model to recognize first names and their variants consistently.
- How to set it up:
- In your collection's Document Understanding settings, go to Entities > Create Entity
- Define an entity type like "FirstName"
- Upload training documents and annotate instances of first names (including variants like "john" and "jon")
- Train the model, then deploy it to your collection
Once deployed, Discovery will automatically recognize these entities across your documents, making searches for name variants more reliable.
内容的提问来源于stack exchange,提问作者ShwetaJ

