如何将CSV文章分词后的词语带句子ID存入新CSV文件?
Hey there! Let's get your script sorted to generate that output.csv exactly how you want it. I'll walk you through the fixes and improvements step by step:
First, fix the small snags in your current code
Your existing code has a couple of minor issues:
- You used
tokenizer.tokenizebut didn't definetokenizer— we'll use NLTK's built-insent_tokenizedirectly - Your while loop only processes the first row of your CSV; we'll adjust it to handle all articles in your file
Full working code
Here's the revised script that does exactly what you need:
import pandas as pd from nltk.tokenize import sent_tokenize, word_tokenize # First, download NLTK's required tokenization resources (run this once) # nltk.download('punkt') # Read your input CSV (remove nrows=10 if you want to process all rows) df = pd.read_csv(r"D:\data.csv") # Initialize variables to track sentence numbers and store our output data sentence_number = 0 output_records = [] # Loop through every article in your CSV for article in df['articles']: # Split the article into individual sentences sentences = sent_tokenize(article) for sentence in sentences: sentence_number += 1 # Split the sentence into words (including punctuation) words = word_tokenize(sentence) # Join the words into a single space-separated string words_string = ' '.join(words) # Add this sentence's data to our output list output_records.append({ 'Sentence No': sentence_number, 'Word': words_string }) # Convert our list of records into a DataFrame output_df = pd.DataFrame(output_records) # Save the DataFrame to output.csv (no extra index column) output_df.to_csv(r"D:\output.csv", index=False)
What this code does:
- Handles all articles: The
for article in df['articles']loop goes through every entry in yourarticlescolumn, not just the first one. - Proper tokenization: Uses NLTK's reliable
sent_tokenizeandword_tokenizeto split text correctly (including punctuation like quotes and periods). - Matches your desired format: Joins each sentence's words into a single string, then pairs it with the incrementing sentence number.
- Clean CSV output: The
index=Falseargument ensures we don't add an extra unnecessary index column tooutput.csv.
Quick note:
If you get an error about missing NLTK resources, uncomment the nltk.download('punkt') line and run it once — this downloads the data needed for tokenization.
When you run this with your sample article, you'll get exactly the output you showed:
Sentence No | Word
1 | The ultimate productivity hack is saying no .
2 | Not doing something will always be faster than doing it .
3 | This statement reminds me of the old computer programming saying , “ Remember that there is no code faster than no code . ”
内容的提问来源于stack exchange,提问作者Abu Toha Md Faisal

