基于Pandas Series生成词云遇ValueError报错的技术求助
Let's break down why your code is failing and how to fix it:
The Root Cause
When you convert each token list to a string (w=str(w)), you turn something like [laptop, sit, star] into "[laptop, sit, star]". Using " ".join(w) on this string splits it into individual characters (like [, l, a, etc.), which aren't valid words. That's why WordCloud can't detect any usable terms, throwing the "got 0 words" error.
Corrected Approach
Step 1: Properly Build Your Text Corpus
Initialize an empty string, then iterate through each list of tokens in tokens_lemma, joining the actual token elements into space-separated strings:
# Start with an empty string to accumulate all words comment_words = "" # Loop through each list of lemmatized tokens for token_list in tokens_lemma: # Join the tokens in the list into a single string with spaces comment_words += " ".join(token_list) + " "
Step 2: Generate the WordCloud
Use your existing plotting code with the correctly built comment_words variable:
wordcloud = WordCloud(width=800, height=800, background_color='white', min_font_size=10).generate(comment_words) # Plot the WordCloud image plt.figure(figsize=(8, 8), facecolor=None) plt.imshow(wordcloud) plt.axis("off") plt.tight_layout(pad=0) plt.show()
Optional: Skip Empty Lists
If some entries in tokens_lemma are empty lists, you can skip them to avoid adding unnecessary spaces:
comment_words = "" for token_list in tokens_lemma: # Only process non-empty token lists if token_list: comment_words += " ".join(token_list) + " "
This will ensure comment_words contains valid, space-separated words that WordCloud can process correctly, resolving the ValueError.
内容的提问来源于stack exchange,提问作者Win Wongsawatdichart

