You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Notebook中提取单列做NLP时的行号干扰问题

Fixing Row Numbers Mixed into User Reviews Column for NLP Processing

Got it, let's work through this problem—dealing with row numbers sneaking into your "User Reviews" column is super common when handling unstructured data, especially if your source data wasn't cleaned properly before importing into your notebook. Here's how to tackle it step by step:

Step 1: Diagnose the Root Cause

First, figure out why those row numbers are there:

  • Case 1: Bad data import: If you loaded your data (e.g., CSV) without specifying an index column, pandas might have treated the original row numbers as part of the "User Reviews" column.
  • Case 2: Data source has embedded row numbers: Some datasets (like exported spreadsheets or scraped data) include row numbers directly in the review text (e.g., 123: Great product! or (456) Love this app).

Step 2: Clean the Column

Fix for Import Issues

If the row numbers came from a bad import, re-read your data with the correct index parameter to separate row numbers from your review text:

import pandas as pd
# Specify the index column to exclude it from your review data
df = pd.read_csv('your_data_file.csv', index_col=0)
# Now extract your reviews safely
text = df.loc[:, "User Reviews"]

Fix for Embedded Row Numbers in Text

If the row numbers are part of the review text itself, use regex to strip them out without touching the legitimate numbers in the reviews. This works for most common row number formats:

import re

def clean_review_text(text):
    # Convert to string to handle NaNs/non-text values
    text_str = str(text).strip()
    # Regex to match leading row numbers (e.g., "123: ", "(456) ", "789. ")
    # This won't affect numbers in the middle of reviews like "I gave it 5 stars"
    cleaned_text = re.sub(r'^\d+[:/.\s]|\(\d+\)\s', '', text_str)
    return cleaned_text if cleaned_text != 'nan' else ''

# Apply the cleaning function to your column
df['Cleaned Reviews'] = df['User Reviews'].apply(clean_review_text)

Pro tip: Test this regex on a few sample rows first to make sure it's not removing anything you want to keep. Adjust the pattern if your row numbers have a weird format (e.g., Row 123: would need ^Row\s\d+:\s).

Step 3: Proceed with NLP (分词 & 高频关键词)

Once your text is clean, you can run your tokenization and keyword stats smoothly. Here's how to do it for both Chinese and English:

For Chinese Text (使用jieba分词)

import jieba
from collections import Counter

# Define a tokenization function with stopword filtering
def tokenize_chinese(text):
    if not text:
        return []
    # Cut text into words
    words = jieba.lcut(text)
    # Basic stopword list (expand this with a full stopword set for better results)
    stop_words = {'的', '了', '是', '我', '你', '就', '都'}
    # Keep only meaningful words (exclude stopwords and single characters)
    return [word for word in words if word not in stop_words and len(word) > 1]

# Apply tokenization
df['Segmented Words'] = df['Cleaned Reviews'].apply(tokenize_chinese)

# Calculate top keywords
all_words = [word for sublist in df['Segmented Words'].tolist() for word in sublist]
top_keywords = Counter(all_words).most_common(10)
print("Top 10 Keywords:", top_keywords)

For English Text (使用NLTK)

import nltk
from nltk.tokenize import word_tokenize
from nltk.corpus import stopwords
from collections import Counter

# Download required NLTK resources (run once)
nltk.download('punkt')
nltk.download('stopwords')

def tokenize_english(text):
    if not text:
        return []
    # Tokenize and convert to lowercase
    tokens = word_tokenize(text.lower())
    # Filter out stopwords and non-alphabetic tokens
    stop_words = set(stopwords.words('english'))
    return [token for token in tokens if token.isalpha() and token not in stop_words]

# Apply tokenization
df['Segmented Words'] = df['Cleaned Reviews'].apply(tokenize_english)

# Calculate top keywords
all_words = [word for sublist in df['Segmented Words'].tolist() for word in sublist]
top_keywords = Counter(all_words).most_common(10)
print("Top 10 Keywords:", top_keywords)

Final Notes

  • Always handle missing values: The clean_review_text function converts NaNs to empty strings to avoid errors during tokenization.
  • Use a comprehensive stopword list: For better results, use a pre-built stopword set (e.g., Chinese stopwords from jieba or extended English stopwords from NLTK).

内容的提问来源于stack exchange,提问作者John Veridan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:42:43