You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将CSV文章分词后的词语带句子ID存入新CSV文件?

Hey there! Let's get your script sorted to generate that output.csv exactly how you want it. I'll walk you through the fixes and improvements step by step:

First, fix the small snags in your current code

Your existing code has a couple of minor issues:

  • You used tokenizer.tokenize but didn't define tokenizer — we'll use NLTK's built-in sent_tokenize directly
  • Your while loop only processes the first row of your CSV; we'll adjust it to handle all articles in your file

Full working code

Here's the revised script that does exactly what you need:

import pandas as pd
from nltk.tokenize import sent_tokenize, word_tokenize

# First, download NLTK's required tokenization resources (run this once)
# nltk.download('punkt')

# Read your input CSV (remove nrows=10 if you want to process all rows)
df = pd.read_csv(r"D:\data.csv")

# Initialize variables to track sentence numbers and store our output data
sentence_number = 0
output_records = []

# Loop through every article in your CSV
for article in df['articles']:
    # Split the article into individual sentences
    sentences = sent_tokenize(article)
    
    for sentence in sentences:
        sentence_number += 1
        # Split the sentence into words (including punctuation)
        words = word_tokenize(sentence)
        # Join the words into a single space-separated string
        words_string = ' '.join(words)
        # Add this sentence's data to our output list
        output_records.append({
            'Sentence No': sentence_number,
            'Word': words_string
        })

# Convert our list of records into a DataFrame
output_df = pd.DataFrame(output_records)

# Save the DataFrame to output.csv (no extra index column)
output_df.to_csv(r"D:\output.csv", index=False)

What this code does:

  1. Handles all articles: The for article in df['articles'] loop goes through every entry in your articles column, not just the first one.
  2. Proper tokenization: Uses NLTK's reliable sent_tokenize and word_tokenize to split text correctly (including punctuation like quotes and periods).
  3. Matches your desired format: Joins each sentence's words into a single string, then pairs it with the incrementing sentence number.
  4. Clean CSV output: The index=False argument ensures we don't add an extra unnecessary index column to output.csv.

Quick note:

If you get an error about missing NLTK resources, uncomment the nltk.download('punkt') line and run it once — this downloads the data needed for tokenization.

When you run this with your sample article, you'll get exactly the output you showed:

Sentence No | Word
1 | The ultimate productivity hack is saying no .
2 | Not doing something will always be faster than doing it .
3 | This statement reminds me of the old computer programming saying , “ Remember that there is no code faster than no code . ”

内容的提问来源于stack exchange,提问作者Abu Toha Md Faisal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:05:54