You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Python逐行提取CSV文本关键词并生成新字段?

Fix: Extract Keywords Per Row from CSV Column

Hey there! The problem with your current code is that you’re concatenating all the text from the some_text column into one big block first—so you’re getting keywords for the entire dataset instead of each individual row. Let’s fix that by processing each row separately.

Modified Code (Per-Row Keyword Extraction)

import pandas as pd
import nltk
from nltk.tokenize import word_tokenize
from nltk.corpus import stopwords

# Download required NLTK resources (run once, then comment out)
nltk.download('punkt')
nltk.download('stopwords')

# Load your CSV data
df = pd.read_csv('test-data.csv')

# Use sets for faster lookup performance
stop_words = set(stopwords.words('english'))
punctuations = {'(',')',';',':','[',']',',','!','?'}

def extract_row_keywords(text):
    # Step 1: Convert text to lowercase for consistent filtering
    text_lower = text.lower()
    # Step 2: Split text into individual tokens
    tokens = word_tokenize(text_lower)
    # Step 3: Filter out stopwords and punctuation
    filtered_tokens = [word for word in tokens if word not in stop_words and word not in punctuations]
    # Step 4: Remove duplicate keywords while keeping their original order
    unique_keywords = list(dict.fromkeys(filtered_tokens))
    # Step 5: Join into a comma-separated string
    return ','.join(unique_keywords)

# Apply the function to every row in the 'some_text' column
df['new_keyword_field'] = df['some_text'].apply(extract_row_keywords)

# Output the result as CSV (matches your expected format)
print(df.to_csv(index=False, quotechar='"'))

Key Changes Explained

  • No more text concatenation: We removed str.cat(sep=' ') so we process each row's text independently.
  • Reusable row-processing function: extract_row_keywords() encapsulates all the logic for a single text entry—easy to tweak later if you need to adjust keyword rules.
  • Faster lookups: Converted stop_words and punctuations to sets (instead of lists) because set membership checks are way faster, especially with large datasets.
  • Deduplicated keywords: Used dict.fromkeys() to keep only unique keywords while preserving their order in the original text (works in Python 3.7+ since dicts maintain insertion order).
  • NLTK resource setup: Added nltk.download() calls to ensure you have the tokenizer and stopword list required for processing (you can comment these out after the first run).

Expected Output

When you run the code, you’ll get exactly the CSV format you wanted:

id,some_text,new_keyword_field
1,"What is the meaning of the word Himalaya?",meaning,word,himalaya
2,"Palindrome is a word, phrase, or sequence that reads the same backward as forward",palindrome,word,phrase,sequence,reads,backward,forward

内容的提问来源于stack exchange,提问作者sharp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:53:26