You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BERTopic报TypeError:确保可迭代对象仅含字符串的解决方法

解决BERTopic中TypeError: "Make sure that the iterable only contains strings"问题

问题场景

作为Python新手,尝试用BERTopic结合PyLDAvis可视化主题建模结果并与LDA对比时,运行代码触发如下TypeError:

TypeError: Make sure that the iterable only contains strings.

报错核心代码片段

import pyLDAvis
import numpy as np
from bertopic import BERTopic

# Train Model
bert_model = BERTopic(verbose=True, calculate_probabilities=True)
topics, probs = bert_model.fit_transform(data_words)  # 触发报错的行

# 后续PyLDAvis准备代码略

完整错误栈

/Library/Frameworks/Python.framework/Versions/3.11/lib/python3.11/site-packages/tqdm/auto.py:21: TqdmWarning: IProgress not found. Please update jupyter and ipywidgets. See https://ipywidgets.readthedocs.io/en/stable/user_install.html
  from .autonotebook import tqdm as notebook_tqdm
---------------------------------------------------------------------------
TypeError                                 Traceback (most recent call last)
Cell In[9], line 4
      1 from bertopic import BERTopic
      3 bert_model = BERTopic()
----> 4 topics, probs = bert_model.fit_transform(data_words)
      6 bert_model.get_topic_freq()

File /Library/Frameworks/Python.framework/Versions/3.11/lib/python3.11/site-packages/bertopic/_bertopic.py:373, in BERTopic.fit_transform(self, documents, embeddings, images, y)
    325 """ Fit the models on a collection of documents, generate topics,
    326 and return the probabilities and topic per document.
    327 
   (...)
    370 ```
    371 """
    372 if documents is not None:
--> 373     check_documents_type(documents)
    374     check_embeddings_shape(embeddings, documents)
    376 doc_ids = range(len(documents)) if documents is not None else range(len(images))

File /Library/Frameworks/Python.framework/Versions/3.11/lib/python3.11/site-packages/bertopic/_utils.py:43, in check_documents_type(documents)
     41 elif isinstance(documents, Iterable) and not isinstance(documents, str):
     42     if not any([isinstance(doc, str) for doc in documents]):
--> 43         raise TypeError("Make sure that the iterable only contains strings.")
     44 else:
     45     raise TypeError("Make sure that the documents variable is an iterable containing strings only.")

TypeError: Make sure that the iterable only contains strings.

数据集与data_words生成逻辑

数据集为JSON格式,每个条目包含tokens字段(拆分后的单词列表):

{
    "TFU_1881_00102": {
        "magazine": "edited out",
        "country": "United Kingdom",
        "year": "1881",
        "tokens": ["word1", "word2"],
        "bigramFreqs": {"word1 word2": 1},
        "tokenFreqs": {"word1": 1, "word2": 1}
    },
    "TFU_1881_00103": {
        "magazine": "edited out",
        "country": "United Kingdom",
        "year": "1881",
        "tokens": ["word3", "word4"],
        "bigramFreqs": {"word3 word4": 1},
        "tokenFreqs": {"word3": 1, "word4": 1}
    }
}

生成data_words的代码:

with open("Data/5_json/output_final.json", "r") as file:
    data = json.load(file)

data_words = []
counter = 0
for key in data:
    counter += 1
    sub_list = data[key]["tokens"]
    data_words.append(sub_list)  # 每个元素是单词列表,而非字符串
print(counter)

错误原因

BERTopic的fit_transform方法要求输入的documents参数是由字符串组成的可迭代对象(每个元素对应一篇文档的完整文本),但当前data_words是列表的列表(每个元素是拆分后的单词列表),不符合输入格式要求,触发类型检查错误。

解决方法

将每个文档的tokens列表拼接成完整字符串,生成符合要求的输入数据:

修正代码(基于原有data_words转换)

# 转换现有data_words为字符串列表
flat_data_words = []
for list_of_strings in data_words:
    sentence = ' '.join(list_of_strings)
    flat_data_words.append(sentence)

# 使用转换后的数据集训练BERTopic
bert_model = BERTopic(verbose=True, calculate_probabilities=True)
topics, probs = bert_model.fit_transform(flat_data_words)

简化版(读取数据时直接转换)

with open("Data/5_json/output_final.json", "r") as file:
    data = json.load(file)

flat_data_words = []
counter = 0
for key in data:
    counter += 1
    # 直接将tokens拼接为字符串
    doc_text = ' '.join(data[key]["tokens"])
    flat_data_words.append(doc_text)
print(counter)

# 后续训练代码不变
bert_model = BERTopic(verbose=True, calculate_probabilities=True)
topics, probs = bert_model.fit_transform(flat_data_words)

内容的提问来源于stack exchange,提问作者Dominik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 20:30:39