Elasticsearch瑞典分析器关键词识别机制及默认列表疑问
Let's break down your questions clearly:
How does the analyzer identify keywords?
In the provided reimplementation of the Swedish analyzer, keyword recognition is handled by the swedish_keywords filter (using the keyword_marker type). Here's the breakdown:
- This filter marks specific words you define as "keywords" during the analysis pipeline.
- Any marked keyword will not be modified by subsequent stemmers (like the
swedish_stemmerin the example). For instance, in the sample config, the wordexempelstays exactly as it is, instead of being stemmed to its root form.
Does Elasticsearch have a predefined list of Swedish keywords?
Short answer: No, there’s no built-in predefined Swedish keyword list for the keyword_marker filter.
By default, if you don’t manually specify values in the settings.analysis.filter.swedish_keywords.keywords field, this filter won’t do anything—there are no hidden default keywords it uses. You have to explicitly define the words you want to protect from stemming (or other text transformations) yourself.
To put it simply: You must define your own keywords for this filter; there’s no pre-existing, ready-to-use Swedish keyword list you can call on out of the box.
内容的提问来源于stack exchange,提问作者Sahand

