使用MiniBatchKMeans的增量式BERTopic处理推特数据集时,半数数据无主题标签的问题求助
使用MiniBatchKMeans的增量式BERTopic处理推特数据集时,半数数据无主题标签的问题求助
我正在尝试给一个包含约5000万条推特的数据集做主题建模。但因为嵌入向量的问题,这么大的数据集即使在128GB内存的机器上也装不下,所以我实现了一个增量式的BERTopic版本,代码如下:
from bertopic.vectorizers import OnlineCountVectorizer from bertopic.vectorizers import ClassTfidfTransformer from sklearn.cluster import MiniBatchKMeans from sklearn.decomposition import IncrementalPCA import numpy as np from tqdm import tqdm import pandas as pd class SafeIncrementalPCA(IncrementalPCA): def partial_fit(self, X, y=None): # Ensure the input is contiguous and in float64 X = np.ascontiguousarray(X, dtype=np.float64) return super().partial_fit(X, y) def transform(self, X): result = super().transform(X) # Force the output to be float64 and contiguous return np.ascontiguousarray(result, dtype=np.float64) vectorizer_model = OnlineCountVectorizer(stop_words="english") ctfidf_model = ClassTfidfTransformer(reduce_frequent_words=True, bm25_weighting=True) umap_model = SafeIncrementalPCA(n_components=100) cluster_model = MiniBatchKMeans(n_clusters=1000, random_state=0) from bertopic import BERTopic topic_model = BERTopic(umap_model=umap_model, hdbscan_model=cluster_model) for docs_delayed, emb_delayed in tqdm(zip(docs_partitions, embeddings_partitions), total=len(docs_partitions)): docs_pdf = docs_delayed.compute() emb_pdf = emb_delayed.compute() docs = docs_pdf["text"].tolist() embeddings = np.vstack(emb_pdf['embeddings'].tolist()) # Partial fit your model (make sure your model supports partial_fit, like many scikit-learn estimators do) topic_model.partial_fit(docs, embeddings)
处理完模型拟合后,我将结果存入SQL数据库,代码如下:
for docs_delayed, emb_delayed in tqdm(zip(docs_partitions, embeddings_partitions), total=len(docs_partitions)): docs_pdf = docs_delayed.compute() emb_pdf = emb_delayed.compute() docs = docs_pdf["text"].tolist() embeddings = np.vstack(emb_pdf['embeddings'].tolist()) # 3) Apply BERTopic on this shard topics, probs = topic_model.transform(docs, embeddings) # Save topics to DataFrame df_topics = pd.DataFrame({ "tweet_id": docs_pdf["id"].tolist(), "topic": topics, "probability": probs }) ## Merge & store in DB docs_pdf["topic"] = df_topics["topic"] docs_pdf["probability"] = df_topics["probability"] docs_pdf.to_sql("tweets", engine, if_exists="append", index=False)
我折腾了好一阵子才写出这个能运行的版本,但现在遇到一个棘手的问题:处理完成后,数据库里有半数数据的主题标签是null。
从理论上来说,MiniBatchKMeans应该不会产生离群点,所有推特都应该被分配到至少一个主题里才对。我特意检查了那些未被分类的推特,它们的内容看起来和已分类的推特没有明显差异,不存在什么特殊的难分类的情况。
希望大家能给我一些建议,看看问题可能出在哪里,以及如何修复这个问题!谢谢!
备注:内容来源于stack exchange,提问作者Matthieu B
相关产品推荐
相关产品推荐

