在ChromaDB中如何为文档关联多元数据字典并实现查询?
ChromaDB 嵌套元数据结构的查询支持与实现方法
一、是否支持查询?
是支持的,但你当前的元数据结构存在明显局限性:
data字段存的是非标准JSON字符串(多个JSON对象直接用逗号拼接,未包裹成数组),ChromaDB只会把它当成普通字符串处理,识别不了内部的实体结构。- 只能基于整个
data字符串做模糊或精确匹配,没法直接针对单个实体的属性(比如entity_group、entity_text)做精准过滤。
二、查询语句编写方法
1. 基于现有结构的字符串匹配查询
如果不想修改数据结构,可以用ChromaDB的元数据过滤语法,对data字段做字符串包含匹配:
示例1:查询包含特定实体文本的文档
collection.query( # 若不需要语义检索,可将query_texts设为空数组 query_texts=["你的查询文本"], where={"data": {"$contains": "Patient"}}, n_results=2 )
示例2:查询包含特定实体组的文档
collection.query( where={"data": {"$contains": "\"entity_group\": \"species\""}}, n_results=2 )
注意这里要转义引号,因为要匹配data字符串里的精确子串。
2. 优化数据结构后的精准查询(推荐)
要实现单个实体属性的精准过滤,建议修改metadatas结构,把data改成实体对象的数组(合法JSON格式),添加数据时的写法如下:
collection.add( documents=["doc1", "doc2"], metadatas=[ { "entities": [ {"entity_text": "Patient", "entity_group": "species", "normalized_id": ["NCBITaxon:9606"], "field": "Exclusion Criteria", "is_negated": "no"}, {"entity_text": "donor", "entity_group": "Subject", "normalized_id": "CUI-less", "field": "Exclusion Criteria", "is_negated": "no"}, {"entity_text": "pregnant lactating", "entity_group": "pregnancy", "normalized_id": "CUI-less", "field": "Exclusion Criteria", "is_negated": "no"} ] }, { "entities": [ {"entity_text": "bilirubin", "entity_group": "drug", "normalized_id": ["mesh:D001663"], "field": "Exclusion Criteria", "is_negated": "no"}, {"entity_text": "serum", "entity_group": "Biological_structure", "normalized_id": "CUI-less", "field": "Exclusion Criteria", "is_negated": "no"}, {"entity_text": "greater than 4 mg / dl", "entity_group": "Lab_value", "normalized_id": "CUI-less", "field": "Exclusion Criteria", "is_negated": "no"}, {"entity_text": "transaminases", "entity_group": "Diagnostic_procedure", "normalized_id": "CUI-less", "field": "Exclusion Criteria", "is_negated": "no"}, {"entity_text": "greater than 5", "entity_group": "Lab_value", "normalized_id": "CUI-less", "field": "Exclusion Criteria", "is_negated": "no"} ] } ], ids=["id1", "id2"] )
修改后就能用ChromaDB的嵌套元数据过滤语法,针对实体属性做精准查询:
示例1:查询包含entity_group为species的文档
collection.query( where={"entities": {"$elemMatch": {"entity_group": "species"}}}, n_results=2 )
示例2:查询包含entity_text为bilirubin且is_negated为no的文档
collection.query( where={"entities": {"$elemMatch": {"entity_text": "bilirubin", "is_negated": "no"}}}, n_results=2 )
内容的提问来源于stack exchange,提问作者Majd Abdallah
相关产品推荐
相关产品推荐

