BERTopic处理17万条数据出现BrokenProcessPool错误如何解决
我使用BERTopic对174,827条数据执行主题建模,运行代码如下:
from bertopic import BERTopic topic_model = BERTopic(language="english", calculate_probabilities=False, verbose=True) topics, probs = topic_model.fit_transform(docs)
运行后返回如下报错信息:
Batches: 100% 5464/5464 [02:11<00:00, 90.71it/s] 2021-11-22 09:36:23,059 - BERTopic - Transformed documents to Embeddings 2021-11-22 09:43:58,215 - BERTopic - Reduced dimensionality with UMAP --------------------------------------------------------------------------- _RemoteTraceback Traceback (most recent call last) _RemoteTraceback: """ Traceback (most recent call last): File "/usr/local/lib/python3.7/dist-packages/joblib/externals/loky/process_executor.py", line 407, in _process_worker call_item = call_queue.get(block=True, timeout=timeout) File "/usr/lib/python3.7/multiprocessing/queues.py", line 113, in get return _ForkingPickler.loads(res) File "sklearn/neighbors/_binary_tree.pxi", line 1057, in sklearn.neighbors._kd_tree.BinaryTree.__setstate__ File "sklearn/neighbors/_binary_tree.pxi", line 999, in sklearn.neighbors._kd_tree.BinaryTree._update_memviews File "stringsource", line 658, in View.MemoryView.memoryview_cwrapper File "stringsource", line 349, in View.MemoryView.memoryview.__cinit__ ValueError: buffer source array is read-only """ The above exception was the direct cause of the following exception: BrokenProcessPool Traceback (most recent call last) <ipython-input-9-ab3893bf488b> in <module>() 2 3 topic_model = BERTopic(language="english", calculate_probabilities=False, verbose=True) ----> 4 topics, probs = topic_model.fit_transform(docs) 10 frames hdbscan/_hdbscan_boruvka.pyx in hdbscan._hdbscan_boruvka.KDTreeBoruvkaAlgorithm.__init__() hdbscan/_hdbscan_boruvka.pyx in hdbscan._hdbscan_boruvka.KDTreeBoruvkaAlgorithm._compute_bounds() /usr/lib/python3.7/concurrent/futures/_base.py in __get_result(self) 382 def __get_result(self): 383 if self._exception: --> 384 raise self._exception 385 else: 386 return self._result BrokenProcessPool: A task has failed to un-serialize. Please ensure that the arguments of the function are all picklable.
相同代码处理约50,000条数据时可正常运行,当前使用搭载GPU的Google Colab环境,请问该如何解决该报错,实现大数据量下的正常主题建模?
报错原因
该问题是旧版本依赖库的已知兼容性问题:大数据量下UMAP输出的降维后数组会被标记为只读,HDBSCAN调用多进程做聚类时,进程间序列化传输该数组触发反序列化失败。小数据量时HDBSCAN默认走单进程逻辑,不会触发该冲突。
解决方法
方法1:自定义HDBSCAN禁用多进程(最稳定,无需升级依赖)
直接给BERTopic传入禁用多进程的HDBSCAN实例,避免进程间序列化步骤:
from bertopic import BERTopic from umap import UMAP from hdbscan import HDBSCAN # 自定义HDBSCAN,核心参数core_dist_n_jobs设为1禁用多进程 hdbscan_model = HDBSCAN( min_cluster_size=10, metric='euclidean', cluster_selection_method='eom', prediction_data=True, core_dist_n_jobs=1 ) # 可选:自定义UMAP关闭低内存模式,降低数组只读概率 umap_model = UMAP( n_neighbors=15, n_components=5, min_dist=0.0, metric='cosine', low_memory=False, random_state=42 ) topic_model = BERTopic( language="english", calculate_probabilities=False, verbose=True, umap_model=umap_model, hdbscan_model=hdbscan_model ) topics, probs = topic_model.fit_transform(docs)
方法2:升级依赖库
该只读数组bug已在高版本scikit-learn、HDBSCAN中修复,Colab中执行如下命令升级依赖,重启运行时后再执行原代码即可:
!pip install --upgrade bertopic hdbscan scikit-learn umap-learn
方法3:传入预计算的可写embedding
自行提前计算文档embedding,拷贝一份得到可写数组后再传入BERTopic,也可避免该问题:
from sentence_transformers import SentenceTransformer # 自行计算embedding embedding_model = SentenceTransformer("all-MiniLM-L6-v2") embeddings = embedding_model.encode(docs, show_progress_bar=True) # 拷贝数组得到可写版本 embeddings = embeddings.copy() topic_model = BERTopic(language="english", calculate_probabilities=False, verbose=True) topics, probs = topic_model.fit_transform(docs, embeddings=embeddings)
内容的提问来源于stack exchange,提问作者AneesBaqir
相关产品推荐
相关产品推荐

