Python Notebook中提取单列做NLP时的行号干扰问题
Got it, let's work through this problem—dealing with row numbers sneaking into your "User Reviews" column is super common when handling unstructured data, especially if your source data wasn't cleaned properly before importing into your notebook. Here's how to tackle it step by step:
Step 1: Diagnose the Root Cause
First, figure out why those row numbers are there:
- Case 1: Bad data import: If you loaded your data (e.g., CSV) without specifying an index column, pandas might have treated the original row numbers as part of the "User Reviews" column.
- Case 2: Data source has embedded row numbers: Some datasets (like exported spreadsheets or scraped data) include row numbers directly in the review text (e.g.,
123: Great product!or(456) Love this app).
Step 2: Clean the Column
Fix for Import Issues
If the row numbers came from a bad import, re-read your data with the correct index parameter to separate row numbers from your review text:
import pandas as pd # Specify the index column to exclude it from your review data df = pd.read_csv('your_data_file.csv', index_col=0) # Now extract your reviews safely text = df.loc[:, "User Reviews"]
Fix for Embedded Row Numbers in Text
If the row numbers are part of the review text itself, use regex to strip them out without touching the legitimate numbers in the reviews. This works for most common row number formats:
import re def clean_review_text(text): # Convert to string to handle NaNs/non-text values text_str = str(text).strip() # Regex to match leading row numbers (e.g., "123: ", "(456) ", "789. ") # This won't affect numbers in the middle of reviews like "I gave it 5 stars" cleaned_text = re.sub(r'^\d+[:/.\s]|\(\d+\)\s', '', text_str) return cleaned_text if cleaned_text != 'nan' else '' # Apply the cleaning function to your column df['Cleaned Reviews'] = df['User Reviews'].apply(clean_review_text)
Pro tip: Test this regex on a few sample rows first to make sure it's not removing anything you want to keep. Adjust the pattern if your row numbers have a weird format (e.g., Row 123: would need ^Row\s\d+:\s).
Step 3: Proceed with NLP (分词 & 高频关键词)
Once your text is clean, you can run your tokenization and keyword stats smoothly. Here's how to do it for both Chinese and English:
For Chinese Text (使用jieba分词)
import jieba from collections import Counter # Define a tokenization function with stopword filtering def tokenize_chinese(text): if not text: return [] # Cut text into words words = jieba.lcut(text) # Basic stopword list (expand this with a full stopword set for better results) stop_words = {'的', '了', '是', '我', '你', '就', '都'} # Keep only meaningful words (exclude stopwords and single characters) return [word for word in words if word not in stop_words and len(word) > 1] # Apply tokenization df['Segmented Words'] = df['Cleaned Reviews'].apply(tokenize_chinese) # Calculate top keywords all_words = [word for sublist in df['Segmented Words'].tolist() for word in sublist] top_keywords = Counter(all_words).most_common(10) print("Top 10 Keywords:", top_keywords)
For English Text (使用NLTK)
import nltk from nltk.tokenize import word_tokenize from nltk.corpus import stopwords from collections import Counter # Download required NLTK resources (run once) nltk.download('punkt') nltk.download('stopwords') def tokenize_english(text): if not text: return [] # Tokenize and convert to lowercase tokens = word_tokenize(text.lower()) # Filter out stopwords and non-alphabetic tokens stop_words = set(stopwords.words('english')) return [token for token in tokens if token.isalpha() and token not in stop_words] # Apply tokenization df['Segmented Words'] = df['Cleaned Reviews'].apply(tokenize_english) # Calculate top keywords all_words = [word for sublist in df['Segmented Words'].tolist() for word in sublist] top_keywords = Counter(all_words).most_common(10) print("Top 10 Keywords:", top_keywords)
Final Notes
- Always handle missing values: The
clean_review_textfunction converts NaNs to empty strings to avoid errors during tokenization. - Use a comprehensive stopword list: For better results, use a pre-built stopword set (e.g., Chinese stopwords from
jiebaor extended English stopwords from NLTK).
内容的提问来源于stack exchange,提问作者John Veridan

