BERTopic报TypeError:确保可迭代对象仅含字符串的解决方法
解决BERTopic中TypeError: "Make sure that the iterable only contains strings"问题
问题场景
作为Python新手,尝试用BERTopic结合PyLDAvis可视化主题建模结果并与LDA对比时,运行代码触发如下TypeError:
TypeError: Make sure that the iterable only contains strings.
报错核心代码片段
import pyLDAvis import numpy as np from bertopic import BERTopic # Train Model bert_model = BERTopic(verbose=True, calculate_probabilities=True) topics, probs = bert_model.fit_transform(data_words) # 触发报错的行 # 后续PyLDAvis准备代码略
完整错误栈
/Library/Frameworks/Python.framework/Versions/3.11/lib/python3.11/site-packages/tqdm/auto.py:21: TqdmWarning: IProgress not found. Please update jupyter and ipywidgets. See https://ipywidgets.readthedocs.io/en/stable/user_install.html from .autonotebook import tqdm as notebook_tqdm --------------------------------------------------------------------------- TypeError Traceback (most recent call last) Cell In[9], line 4 1 from bertopic import BERTopic 3 bert_model = BERTopic() ----> 4 topics, probs = bert_model.fit_transform(data_words) 6 bert_model.get_topic_freq() File /Library/Frameworks/Python.framework/Versions/3.11/lib/python3.11/site-packages/bertopic/_bertopic.py:373, in BERTopic.fit_transform(self, documents, embeddings, images, y) 325 """ Fit the models on a collection of documents, generate topics, 326 and return the probabilities and topic per document. 327 (...) 370 ``` 371 """ 372 if documents is not None: --> 373 check_documents_type(documents) 374 check_embeddings_shape(embeddings, documents) 376 doc_ids = range(len(documents)) if documents is not None else range(len(images)) File /Library/Frameworks/Python.framework/Versions/3.11/lib/python3.11/site-packages/bertopic/_utils.py:43, in check_documents_type(documents) 41 elif isinstance(documents, Iterable) and not isinstance(documents, str): 42 if not any([isinstance(doc, str) for doc in documents]): --> 43 raise TypeError("Make sure that the iterable only contains strings.") 44 else: 45 raise TypeError("Make sure that the documents variable is an iterable containing strings only.") TypeError: Make sure that the iterable only contains strings.
数据集与data_words生成逻辑
数据集为JSON格式,每个条目包含tokens字段(拆分后的单词列表):
{ "TFU_1881_00102": { "magazine": "edited out", "country": "United Kingdom", "year": "1881", "tokens": ["word1", "word2"], "bigramFreqs": {"word1 word2": 1}, "tokenFreqs": {"word1": 1, "word2": 1} }, "TFU_1881_00103": { "magazine": "edited out", "country": "United Kingdom", "year": "1881", "tokens": ["word3", "word4"], "bigramFreqs": {"word3 word4": 1}, "tokenFreqs": {"word3": 1, "word4": 1} } }
生成data_words的代码:
with open("Data/5_json/output_final.json", "r") as file: data = json.load(file) data_words = [] counter = 0 for key in data: counter += 1 sub_list = data[key]["tokens"] data_words.append(sub_list) # 每个元素是单词列表,而非字符串 print(counter)
错误原因
BERTopic的fit_transform方法要求输入的documents参数是由字符串组成的可迭代对象(每个元素对应一篇文档的完整文本),但当前data_words是列表的列表(每个元素是拆分后的单词列表),不符合输入格式要求,触发类型检查错误。
解决方法
将每个文档的tokens列表拼接成完整字符串,生成符合要求的输入数据:
修正代码(基于原有data_words转换)
# 转换现有data_words为字符串列表 flat_data_words = [] for list_of_strings in data_words: sentence = ' '.join(list_of_strings) flat_data_words.append(sentence) # 使用转换后的数据集训练BERTopic bert_model = BERTopic(verbose=True, calculate_probabilities=True) topics, probs = bert_model.fit_transform(flat_data_words)
简化版(读取数据时直接转换)
with open("Data/5_json/output_final.json", "r") as file: data = json.load(file) flat_data_words = [] counter = 0 for key in data: counter += 1 # 直接将tokens拼接为字符串 doc_text = ' '.join(data[key]["tokens"]) flat_data_words.append(doc_text) print(counter) # 后续训练代码不变 bert_model = BERTopic(verbose=True, calculate_probabilities=True) topics, probs = bert_model.fit_transform(flat_data_words)
内容的提问来源于stack exchange,提问作者Dominik
相关产品推荐
相关产品推荐

