Couchbase Python SDK 3.2.x内存泄漏及关闭链路追踪方法求助
解决方案
1 正确禁用链路追踪的配置
Couchbase Python SDK 3.x 版本中,enable_tracing 连接串参数已失效,需要通过 ClusterTracingOptions 显式将所有追踪阈值、队列大小设置为0来完全关闭默认的阈值追踪和孤儿请求追踪,传入None会自动 fallback 到默认配置,无法生效:
from couchbase.cluster import Cluster, ClusterOptions from couchbase.auth import PasswordAuthenticator from couchbase.options import ClusterTracingOptions # 完全关闭所有追踪的配置 tracing_opts = ClusterTracingOptions( tracing_threshold_kv=0, tracing_threshold_view=0, tracing_threshold_query=0, tracing_threshold_search=0, tracing_threshold_analytics=0, tracing_threshold_queue_size=0, tracing_threshold_queue_flush_interval=0, tracing_orphaned_queue_size=0, tracing_orphaned_queue_flush_interval=0 ) # 全局初始化Cluster,不要每次执行任务都新建 cluster = Cluster( COUCH_HOST, ClusterOptions( PasswordAuthenticator(COUCH_USER, COUCH_PASS), tracing_options=tracing_opts ) ) # 提前初始化bucket和collection复用 bucket = cluster.bucket(BUCKET_NAME) collection = bucket.default_collection()
2 规避内存泄漏的额外注意事项
- Cluster 是重量级资源,禁止每次业务函数调用都新建实例,应该全局复用,多进程场景下每个子进程单独初始化一次即可
- 业务执行完成后,需要显式调用
cluster.close()释放底层连接、追踪队列等资源,你的示例代码中没有释放Cluster资源,会直接导致内存持续增长 - 不需要链路追踪的情况下不要传入任何自定义Tracer参数,传入
CouchbaseOtelTracer反而会启动追踪逻辑,增加额外内存开销
3 修改后的函数示例
import time from couchbase.exceptions import PathExistsException, DocumentNotFoundException # Cluster、collection提前全局初始化,不要放在函数内 # 这里省略前面的初始化代码 def multi_proc_insert(a_id, f_path): lines = mapcount(f_path) logger.info(f"Start to process {f_path}\t lines: {lines}") inserted = 0 t1 = time.time() with open(f_path, "r") as fp: while True: e_hash = fp.readline() if not e_hash: logger.info(f"e_hash not found (EOF)") break key = DOC_TYPE + e_hash.strip() try: collection.mutate_in(key, [SD.array_addunique("", a_id)]) except PathExistsException: logger.debug(f"a_id {a_id} already exists on doc {key}") except DocumentNotFoundException: collection.insert(key, [a_id]) except Exception: logger.exception( f"Error processing {f_path}, inserted rows: {inserted} remaining: {lines - inserted}" ) inserted += 1 t2 = time.time() logger.info( f"Finished with {f_path}\t lines: {lines} inserted: {inserted} remaining: {lines - inserted} in seconds {t2 - t1}" ) # 所有任务执行完成后显式关闭Cluster # cluster.close()
内容的提问来源于stack exchange,提问作者Emiliano Dalla Verde Marcozzi
相关产品推荐
相关产品推荐

