在Solr 7.3(SolrCloud)中使用NER的配置疑问及文档补充需求
我来帮你梳理完整的Solr 7.3(SolrCloud)环境下OpenNLP命名实体提取(NER)的配置步骤,补全你需要的细节:
完整配置步骤与细节补充
1. 前置准备
- 先下载对应OpenNLP模型文件:常用的有
en-ner-person.bin、en-ner-location.bin、en-ner-organization.bin、en-ner-date.bin等,把它们放到Solr配置集的conf/opennlp目录下(SolrCloud环境要确保这个目录同步到ZooKeeper)。 - 确认
solr-analysis-extras模块已启用:Solr 7.3发行版默认包含该模块,但自定义部署时要保证solr-analysis-extras-7.3.0.jar在Solr的classpath中。
2. solrconfig.xml 完整配置
你已经写了部分配置,以下是补全后的完整更新处理器链配置:
<updateRequestProcessorChain name="multiple-extract"> <!-- OpenNLP NER提取处理器核心配置 --> <processor class="solr.OpenNLPExtractNamedEntitiesUpdateProcessorFactory"> <!-- 配置多个实体模型文件,每个模型对应一种实体类型 --> <str name="modelFile">opennlp/en-ner-person.bin</str> <str name="modelFile">opennlp/en-ner-location.bin</str> <str name="modelFile">opennlp/en-ner-organization.bin</str> <!-- 指定要从中提取实体的源字段(比如你的文档内容字段) --> <str name="inputField">content</str> <!-- 提取后实体的存储字段前缀,最终会生成如entities_person的字段 --> <str name="outputField">entities_</str> <!-- 可选:开启实体去重,避免同一实体重复存储 --> <bool name="distinct">true</bool> <!-- 可选:保留实体原始大小写,比如"New York"不会转为小写 --> <bool name="preserveCase">true</bool> </processor> <!-- 必须添加默认更新处理器收尾,确保文档正常写入索引 --> <processor class="solr.RunUpdateProcessorFactory"/> </updateRequestProcessorChain>
3. 字段定义(schema.xml/managed-schema)
需要为提取出的实体字段定义类型和字段,示例如下:
<!-- 定义多值字符串类型,适合存储实体集合 --> <fieldType name="strings" class="solr.StrField" multiValued="true" indexed="true" stored="true"/> <!-- 对应NER提取出的各类实体字段 --> <field name="entities_person" type="strings"/> <field name="entities_location" type="strings"/> <field name="entities_organization" type="strings"/>
如果使用managed-schema,也可以通过API添加字段:
curl -X POST -H "Content-Type: application/json" http://<solr-host>:<port>/solr/<collection-name>/schema -d '{ "add-field": [ {"name":"entities_person", "type":"strings", "multiValued":true, "indexed":true, "stored":true}, {"name":"entities_location", "type":"strings", "multiValued":true, "indexed":true, "stored":true}, {"name":"entities_organization", "type":"strings", "multiValued":true, "indexed":true, "stored":true} ] }'
4. SolrCloud环境专属注意事项
- 同步配置到ZooKeeper:用
bin/solr zk upconfig命令把本地配置上传到ZK,确保所有节点获取最新配置:
bin/solr zk upconfig -n <your-configset-name> -d ./path/to/your/config -z <zk-host>:<zk-port>
- 上传模型文件到ZK:用
bin/solr zk putfile命令把模型文件同步到ZK的配置集目录:
bin/solr zk putfile /configsets/<your-configset>/conf/opennlp/en-ner-person.bin ./opennlp/en-ner-person.bin -z <zk-host>:<zk-port>
- 重启所有Solr节点,确保配置和模型生效。
5. 测试配置是否生效
提交测试文档验证:
curl -X POST -H "Content-Type: application/json" http://<solr-host>:<port>/solr/<collection-name>/update?update.chain=multiple-extract -d '[ {"id": "1", "content": "Apple was founded by Steve Jobs in Cupertino, California in 1976."} ]'
查询该文档后,应该能看到entities_person字段包含["Steve Jobs"],entities_organization包含["Apple"],entities_location包含["Cupertino", "California"]。
内容的提问来源于stack exchange,提问作者Alexey Ponomarenko
相关产品推荐
相关产品推荐

