You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用MiniBatchKMeans的增量式BERTopic处理推特数据集时,半数数据无主题标签的问题求助

使用MiniBatchKMeans的增量式BERTopic处理推特数据集时,半数数据无主题标签的问题求助

我正在尝试给一个包含约5000万条推特的数据集做主题建模。但因为嵌入向量的问题,这么大的数据集即使在128GB内存的机器上也装不下,所以我实现了一个增量式的BERTopic版本,代码如下:

from bertopic.vectorizers import OnlineCountVectorizer
from bertopic.vectorizers import ClassTfidfTransformer
from sklearn.cluster import MiniBatchKMeans
from sklearn.decomposition import IncrementalPCA
import numpy as np
from tqdm import tqdm
import pandas as pd


class SafeIncrementalPCA(IncrementalPCA):
    def partial_fit(self, X, y=None):
        # Ensure the input is contiguous and in float64
        X = np.ascontiguousarray(X, dtype=np.float64)
        return super().partial_fit(X, y)
    
    def transform(self, X):
        result = super().transform(X)
        # Force the output to be float64 and contiguous
        return np.ascontiguousarray(result, dtype=np.float64)


vectorizer_model = OnlineCountVectorizer(stop_words="english")
ctfidf_model = ClassTfidfTransformer(reduce_frequent_words=True, bm25_weighting=True)
umap_model = SafeIncrementalPCA(n_components=100)
cluster_model = MiniBatchKMeans(n_clusters=1000, random_state=0)

from bertopic import BERTopic

topic_model = BERTopic(umap_model=umap_model,
                       hdbscan_model=cluster_model)

for docs_delayed, emb_delayed in tqdm(zip(docs_partitions, embeddings_partitions), total=len(docs_partitions)):
    docs_pdf = docs_delayed.compute()
    emb_pdf = emb_delayed.compute()

    docs = docs_pdf["text"].tolist()
    embeddings = np.vstack(emb_pdf['embeddings'].tolist())
    
    # Partial fit your model (make sure your model supports partial_fit, like many scikit-learn estimators do)
    topic_model.partial_fit(docs, embeddings)

处理完模型拟合后,我将结果存入SQL数据库,代码如下:

for docs_delayed, emb_delayed in tqdm(zip(docs_partitions, embeddings_partitions), total=len(docs_partitions)):
    docs_pdf = docs_delayed.compute()
    emb_pdf = emb_delayed.compute()
    docs = docs_pdf["text"].tolist()
    embeddings = np.vstack(emb_pdf['embeddings'].tolist())

    # 3) Apply BERTopic on this shard
    topics, probs = topic_model.transform(docs, embeddings)

    # Save topics to DataFrame
    df_topics = pd.DataFrame({
        "tweet_id": docs_pdf["id"].tolist(),
        "topic": topics,
        "probability": probs
    })

    ## Merge & store in DB
    docs_pdf["topic"] = df_topics["topic"]
    docs_pdf["probability"] = df_topics["probability"]
    docs_pdf.to_sql("tweets", engine, if_exists="append", index=False)

我折腾了好一阵子才写出这个能运行的版本,但现在遇到一个棘手的问题:处理完成后,数据库里有半数数据的主题标签是null。

从理论上来说,MiniBatchKMeans应该不会产生离群点,所有推特都应该被分配到至少一个主题里才对。我特意检查了那些未被分类的推特,它们的内容看起来和已分类的推特没有明显差异,不存在什么特殊的难分类的情况。

希望大家能给我一些建议,看看问题可能出在哪里,以及如何修复这个问题!谢谢!


备注:内容来源于stack exchange,提问作者Matthieu B

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 10:39:36