使用textnets进行文本网络分析时遇ValueError:不允许负维度
ValueError: negative dimensions are not allowed 问题排查与解决
问题场景
完成文本情感分析后,使用textnets库进行文本网络分析,运行t = tn.Textnet(corpus.tokenized())时触发ValueError: negative dimensions are not allowed错误。
完整代码
import textnets as tn tn.params["seed"] = 42 df = pd.read_csv(r"C:\Users\User\Downloads\archive\nltk_split.csv") df.head() df.drop(['Unnamed: 0.2', 'Unnamed: 0.1', 'Unnamed: 0', 'Unnamed: 15', 'Unnamed: 16', 'id', 'favorite_count', 'created_at', 'retweet_count', 'coordinates', 'score', 'neg', 'neu', 'pos', 'compound','created_at_date'], axis=1, inplace=True) df.head() corpus = tn.Corpus.from_df(df, doc_col='text') t = tn.Textnet(corpus.tokenized()) t.plot(label_nodes=True, show_clusters=True)
报错堆栈
--------------------------------------------------------------------------- ValueError Traceback (most recent call last) ~\AppData\Local\Temp\ipykernel_19036\2509364554.py in <module> ----> 1 t = tn.Textnet(corpus.tokenized()) ~\AppData\Roaming\Python\Python39\site-packages\textnets\network.py in __init__(self, data, min_docs, connected, remove_weak_edges, doc_attrs) 353 self._matrix = data 354 elif isinstance(data, (TidyText, pd.DataFrame)): ---> 355 self._matrix = _im_from_tidy_text(data, min_docs) 356 if remove_weak_edges: 357 pairs: pd.Series = self._matrix.stack() ~\AppData\Roaming\Python\Python39\site-packages\textnets\network.py in _im_from_tidy_text(tidy_text, min_docs) 812 .set_index("label") 813 ) ---> 814 im = tt[tt["keep"]].pivot(values="term_weight", columns="term").fillna(0) 815 return IncidenceMatrix(im) 816 E:\anaconda\lib\site-packages\pandas\core\frame.py in pivot(self, index, columns, values) 7883 from pandas.core.reshape.pivot import pivot 7884 -> 7885 return pivot(self, index=index, columns=columns, values=values) 7886 7887 _shared_docs[ E:\anaconda\lib\site-packages\pandas\core\reshape\pivot.py in pivot(data, index, columns, values) 518 else: 519 indexed = data._constructor_sliced(data[values]._values, index=multiindex) ---> 520 return indexed.unstack(columns_listlike) 521 522 E:\anaconda\lib\site-packages\pandas\core\series.py in unstack(self, level, fill_value) 4155 from pandas.core.reshape.reshape import unstack 4156 -> 4157 return unstack(self, level, fill_value) 4158 4159 # ---------------------------------------------------------------------- E:\anaconda\lib\site-packages\pandas\core\reshape\reshape.py in unstack(obj, level, fill_value) 489 if is_1d_only_ea_dtype(obj.dtype): 490 return _unstack_extension_series(obj, level, fill_value) ---> 491 unstacker = _Unstacker( 492 obj.index, level=level, constructor=obj._constructor_expanddim 493 ) E:\anaconda\lib\site-packages\pandas\core\reshape\reshape.py in __init__(self, index, level, constructor) 138 ) 139 ---> 140 self._make_selectors() 141 142 @cache_readonly E:\anaconda\lib\site-packages\pandas\core\reshape\reshape.py in _make_selectors(self) 186 187 selector = self.sorted_labels[-1] + stride * comp_index + self.lift ---> 188 mask = np.zeros(np.prod(self.full_shape), dtype=bool) 189 mask.put(selector, True) 190 ValueError: negative dimensions are not allowed
错误原因
该错误源于pandas的pivot/unstack操作时生成的数组维度为负数,核心诱因包括:
- 输入的文本数据存在大量空值或空白内容,分词后无有效术语
- textnets默认
min_docs=2(术语至少出现在2篇文档中),若所有术语仅出现1次,过滤后无有效数据 - 分词后的tidy文本存在重复的
(文档标签, 术语)组合,导致维度计算异常
解决方法
1. 清理文本数据
先检查并移除空值和空白文本:
# 检查text列空值数量 print(df['text'].isnull().sum()) # 移除空值行 df = df.dropna(subset=['text']) # 移除纯空白文本行 df = df[df['text'].str.strip() != '']
2. 调整min_docs参数
降低术语出现的文档数阈值,比如设为1:
t = tn.Textnet(corpus.tokenized(), min_docs=1)
3. 验证分词结果
确认分词后存在有效术语:
tokenized_corpus = corpus.tokenized() # 查看分词数据前几行 print(tokenized_corpus.head()) # 检查唯一术语数量 print(f"唯一术语数量: {tokenized_corpus['term'].nunique()}")
4. 处理重复条目
若存在重复的(文档标签, 术语)组合,先去重:
tokenized_corpus = tokenized_corpus.drop_duplicates(subset=['label', 'term']) t = tn.Textnet(tokenized_corpus)
内容的提问来源于stack exchange,提问作者Karthik Bhandary
相关产品推荐
相关产品推荐

