You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

高效实现文本分句与清洗的方法(Pandas+NLTK实操)

解决Pandas并行移除停用词时的AttributeError及高效实现方式

错误原因

触发AttributeError: 'int' object has no attribute 'split'的核心问题是你的sentence列里混了整数类型的数据——大概率是原始文本为空/缺失时被误填成了0,或者分句过程中生成了非字符串元素(比如NaN被转成int)。处理前得先把这些无效数据清掉。

高效实现步骤

1. 先清洗sentence列

确保sentence列里的每个元素都是纯字符串列表,过滤掉非列表、列表里的非字符串内容:

import pandas as pd
from nltk.corpus import stopwords
from nltk.tokenize import word_tokenize

# 加载西班牙语停用词,转成集合(查询更快)
spanish_stopwords = set(stopwords.words('spanish'))

# 清洗:保留列表中的有效字符串,非列表转为空列表
df['sentence'] = df['sentence'].apply(
    lambda x: [s for s in x if isinstance(s, str)] if isinstance(x, list) else []
)
# 可选:删掉sentence为空列表的行,避免后续无意义处理
df = df[df['sentence'].str.len() > 0].reset_index(drop=True)

2. 移除停用词的两种高效方式

方式一:常规apply(中小数据集首选)

用列表推导结合apply,代码简洁且性能足够,比手动并行更易维护:

def clean_sentences(sentences):
    cleaned = []
    for sent in sentences:
        # 分词→过滤停用词→重组句子
        tokens = word_tokenize(sent, language='spanish')
        filtered = [tok for tok in tokens if tok.lower() not in spanish_stopwords]
        cleaned.append(' '.join(filtered))
    return cleaned

df['cleaned_sentences'] = df['sentence'].apply(clean_sentences)

方式二:并行加速(大数据集用)

如果数据量特别大,用swifter库自动适配并行逻辑,比自己写multiprocessing省心:

import swifter

df['cleaned_sentences'] = df['sentence'].swifter.apply(clean_sentences)

3. 前置优化:分句时就堵漏洞

在最初的分句步骤里就处理空文本,从源头避免无效数据:

from nltk.tokenize import sent_tokenize

def split_sentences(text):
    # 过滤空字符串和非字符串类型
    if not isinstance(text, str) or text.strip() == '':
        return []
    return sent_tokenize(text, language='spanish')

# 替换你原来的分句代码
df['sentence'] = df['text'].apply(split_sentences)

内容的提问来源于stack exchange,提问作者Luis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 08:40:25