Python新手求助:如何为print输出创建变量用于nltk分词?
Hey there! No worries at all—this is a really common first step when working with text data in Python, and I’ll walk you through it clearly.
First, let’s fix a tiny syntax issue in your original code: your print line is missing a closing parenthesis, which would throw an error. But more importantly, instead of just printing each value, we’ll collect them into a list (a Python variable that holds multiple items) so you can use them for NLTK later.
Step 1: Collect Excel Column Data into a Variable
Here’s how to modify your code to store all the values from column 2 (index 1) in a list:
import xlrd # Make sure your file name includes the extension (like .xls or .xlsx) file_name = "D:/Uber/reviews.xls" workbook = xlrd.open_workbook(file_name) sheet = workbook.sheet_by_index(0) # Create an empty list to hold all your review text review_texts = [] # Loop through each row and add the column 1 value to the list for row in range(sheet.nrows): # Get the value from the current row, column 1 current_text = sheet.cell_value(row, 1) # Add it to our list review_texts.append(current_text) # Now you can check what's in the list (optional) print(review_texts)
Step 2: Use the Variable for NLTK Tokenization
Once you have review_texts (your list of all column values), you can loop through each item and apply NLTK’s tokenizer. Here’s how:
import nltk # You only need to run this once to download the tokenization model nltk.download('punkt') # Create another list to hold the tokenized results tokenized_reviews = [] for text in review_texts: # Tokenize the current text string tokens = nltk.word_tokenize(text) # Add the tokens to our new list tokenized_reviews.append(tokens) # Print out the tokenized reviews to verify for tokens in tokenized_reviews: print(tokens)
Quick Notes:
- Double-check your file path: Ensure
reviewshas the correct extension (.xlsfor older Excel files; if you’re using.xlsx, note that newer versions of xlrd don’t support it—you might need to useopenpyxlinstead if that’s the case). - If your first row is a header (like "Review Text"), you can start the loop at
range(1, sheet.nrows)to skip it.
内容的提问来源于stack exchange,提问作者Thy Quang Lam

