调用fit_transform报错缺少raw_documents参数,该参数指什么?
raw_documents Parameter in TfidfVectorizer.fit_transform() Hey there! Let's clear up this confusion for you quickly.
What exactly is raw_documents?
raw_documents is the required positional parameter for the fit_transform() method of TfidfVectorizer. It refers directly to your collection of text data that you want to process into TF-IDF features. In your case, this is exactly your Reviews list—the restaurant comments you scraped from Yelp!
Why are you seeing this error?
The error pops up because when you called tfidf_vectorizer.fit_transform(), you didn’t pass in any data for it to work with. The method needs that raw text input to learn vocabulary, calculate IDF weights, and convert your reviews into the numerical TF-IDF matrix you’re trying to create.
Fixing your code
You just need to pass your Reviews list as the raw_documents argument. Here’s how your completed code should look (assuming your tokenize_and_stem function is already defined):
from sklearn.feature_extraction.text import TfidfVectorizer # Initialize your vectorizer with your custom parameters tfidf_vectorizer = TfidfVectorizer( max_df=0.8, max_features=200000, min_df=0.2, stop_words='english', use_idf=True, tokenizer=tokenize_and_stem, ngram_range=(1,3) ) # Pass your Reviews list to fit_transform() %time tfidf_matrix = tfidf_vectorizer.fit_transform(Reviews)
A quick extra note
The fit_transform() method combines two steps:
fit(): Learns the vocabulary and IDF values from your raw texttransform(): Converts your raw text into a TF-IDF matrix
Both steps depend on having access to your original review data, which is why raw_documents is non-negotiable here.
内容的提问来源于stack exchange,提问作者user9477395

