高效实现文本分句与清洗的方法(Pandas+NLTK实操)
解决Pandas并行移除停用词时的AttributeError及高效实现方式
错误原因
触发AttributeError: 'int' object has no attribute 'split'的核心问题是你的sentence列里混了整数类型的数据——大概率是原始文本为空/缺失时被误填成了0,或者分句过程中生成了非字符串元素(比如NaN被转成int)。处理前得先把这些无效数据清掉。
高效实现步骤
1. 先清洗sentence列
确保sentence列里的每个元素都是纯字符串列表,过滤掉非列表、列表里的非字符串内容:
import pandas as pd from nltk.corpus import stopwords from nltk.tokenize import word_tokenize # 加载西班牙语停用词,转成集合(查询更快) spanish_stopwords = set(stopwords.words('spanish')) # 清洗:保留列表中的有效字符串,非列表转为空列表 df['sentence'] = df['sentence'].apply( lambda x: [s for s in x if isinstance(s, str)] if isinstance(x, list) else [] ) # 可选:删掉sentence为空列表的行,避免后续无意义处理 df = df[df['sentence'].str.len() > 0].reset_index(drop=True)
2. 移除停用词的两种高效方式
方式一:常规apply(中小数据集首选)
用列表推导结合apply,代码简洁且性能足够,比手动并行更易维护:
def clean_sentences(sentences): cleaned = [] for sent in sentences: # 分词→过滤停用词→重组句子 tokens = word_tokenize(sent, language='spanish') filtered = [tok for tok in tokens if tok.lower() not in spanish_stopwords] cleaned.append(' '.join(filtered)) return cleaned df['cleaned_sentences'] = df['sentence'].apply(clean_sentences)
方式二:并行加速(大数据集用)
如果数据量特别大,用swifter库自动适配并行逻辑,比自己写multiprocessing省心:
import swifter df['cleaned_sentences'] = df['sentence'].swifter.apply(clean_sentences)
3. 前置优化:分句时就堵漏洞
在最初的分句步骤里就处理空文本,从源头避免无效数据:
from nltk.tokenize import sent_tokenize def split_sentences(text): # 过滤空字符串和非字符串类型 if not isinstance(text, str) or text.strip() == '': return [] return sent_tokenize(text, language='spanish') # 替换你原来的分句代码 df['sentence'] = df['text'].apply(split_sentences)
内容的提问来源于stack exchange,提问作者Luis
相关产品推荐
相关产品推荐

