如何基于Python逐行提取CSV文本关键词并生成新字段?
Fix: Extract Keywords Per Row from CSV Column
Hey there! The problem with your current code is that you’re concatenating all the text from the some_text column into one big block first—so you’re getting keywords for the entire dataset instead of each individual row. Let’s fix that by processing each row separately.
Modified Code (Per-Row Keyword Extraction)
import pandas as pd import nltk from nltk.tokenize import word_tokenize from nltk.corpus import stopwords # Download required NLTK resources (run once, then comment out) nltk.download('punkt') nltk.download('stopwords') # Load your CSV data df = pd.read_csv('test-data.csv') # Use sets for faster lookup performance stop_words = set(stopwords.words('english')) punctuations = {'(',')',';',':','[',']',',','!','?'} def extract_row_keywords(text): # Step 1: Convert text to lowercase for consistent filtering text_lower = text.lower() # Step 2: Split text into individual tokens tokens = word_tokenize(text_lower) # Step 3: Filter out stopwords and punctuation filtered_tokens = [word for word in tokens if word not in stop_words and word not in punctuations] # Step 4: Remove duplicate keywords while keeping their original order unique_keywords = list(dict.fromkeys(filtered_tokens)) # Step 5: Join into a comma-separated string return ','.join(unique_keywords) # Apply the function to every row in the 'some_text' column df['new_keyword_field'] = df['some_text'].apply(extract_row_keywords) # Output the result as CSV (matches your expected format) print(df.to_csv(index=False, quotechar='"'))
Key Changes Explained
- No more text concatenation: We removed
str.cat(sep=' ')so we process each row's text independently. - Reusable row-processing function:
extract_row_keywords()encapsulates all the logic for a single text entry—easy to tweak later if you need to adjust keyword rules. - Faster lookups: Converted
stop_wordsandpunctuationsto sets (instead of lists) because set membership checks are way faster, especially with large datasets. - Deduplicated keywords: Used
dict.fromkeys()to keep only unique keywords while preserving their order in the original text (works in Python 3.7+ since dicts maintain insertion order). - NLTK resource setup: Added
nltk.download()calls to ensure you have the tokenizer and stopword list required for processing (you can comment these out after the first run).
Expected Output
When you run the code, you’ll get exactly the CSV format you wanted:
id,some_text,new_keyword_field 1,"What is the meaning of the word Himalaya?",meaning,word,himalaya 2,"Palindrome is a word, phrase, or sequence that reads the same backward as forward",palindrome,word,phrase,sequence,reads,backward,forward
内容的提问来源于stack exchange,提问作者sharp
相关产品推荐
相关产品推荐

