如何在Watson Discovery v2中统计集合内的文档数量?
统计Watson Discovery v2集合文档数量的简便方法
Watson Discovery v2确实移除了v1版本中get_collection接口返回的document_counts字段,不过可以通过以下两种高效方式实现文档计数,避免全量查询的超时问题:
方法1:利用聚合查询快速统计
通过Discovery v2的query接口聚合功能,仅请求文档总数,不返回实际文档内容,响应速度快且无超时风险:
from ibm_watson import DiscoveryV2 from ibm_cloud_sdk_core.authenticators import IAMAuthenticator def get_collection_doc_count(apikey, service_url, project_id, collection_id): authenticator = IAMAuthenticator(apikey) discovery = DiscoveryV2( version='2021-08-30', authenticator=authenticator ) discovery.set_service_url(service_url) # 发起聚合查询,仅统计文档总数 response = discovery.query( project_id=project_id, collection_ids=[collection_id], aggregations=['count()'], return_fields='', # 不返回任何文档字段,减少数据传输 limit=0 # 不返回具体文档结果 ).get_result() # 提取聚合结果中的文档总数 return response['aggregations'][0]['count']
方法2:从集合索引状态获取计数
如果集合近期完成过索引任务,可以通过get_collection接口返回的索引状态元数据间接获取已处理文档数:
from ibm_watson import DiscoveryV2 from ibm_cloud_sdk_core.authenticators import IAMAuthenticator def get_count_from_index_status(apikey, service_url, project_id, collection_id): authenticator = IAMAuthenticator(apikey) discovery = DiscoveryV2( version='2021-08-30', authenticator=authenticator ) discovery.set_service_url(service_url) coll_details = discovery.get_collection( project_id=project_id, collection_id=collection_id ).get_result() # 提取已完成索引的文档数量 return coll_details['index_status']['documents_processed']
注意:该方法返回的是已完成索引的文档数,若有正在处理的文档,数值会和实际可用文档数存在偏差,适合对实时性要求不高的场景。
内容的提问来源于stack exchange,提问作者Giacomo Bartoli
相关产品推荐
相关产品推荐

