使用spaCy获取文本统计数据时如何过滤换行符与无意义短句?
spaCy分句结果无效内容过滤方案
核心过滤规则
- 去除换行符与空白冗余:先清洗分句结果中的换行符、首尾空白,清洗后为空的分句直接丢弃
- 过滤过短短句:调整token计数逻辑,仅统计非标点、非空格的有效token,自行设定最小有效token阈值(一般建议设为2~3),低于阈值的短句直接过滤
- 过滤无意义语句:可通过词性规则校验,要求句子至少包含1个名词/代词 + 1个动词,也可自定义规则过滤全特殊符号、全数字、无意义重复内容等异常语句
修改后实现代码
import spacy import pandas as pd nlp = spacy.load("en_core_web_md") # 可根据业务需求调整阈值 MIN_VALID_TOKENS = 3 def describe_sentence_stats(text): doc = nlp(text) sent_data = [] for sent in doc.sents: # 第一步:清洗换行符与空白 clean_sent = sent.text.strip().replace('\n', ' ').replace('\r', '') if not clean_sent: continue # 第二步:统计有效token(排除标点、空格) valid_tokens = [token for token in sent if not token.is_punct and not token.is_space] if len(valid_tokens) < MIN_VALID_TOKENS: continue # 第三步:校验语句有效性(示例规则:至少包含1个名词/代词 + 1个动词) has_noun = any(token.pos_ in ('NOUN', 'PROPN', 'PRON') for token in valid_tokens) has_verb = any(token.pos_ in ('VERB', 'AUX') for token in valid_tokens) if not (has_noun and has_verb): continue # 符合要求的句子加入结果 sent_data.append({ 'sentence': clean_sent, 'num_words': len(valid_tokens) }) return pd.DataFrame(sent_data) sentences = describe_sentence_stats(example_text)
自定义调整建议
- 如果你的场景允许无动词的短句(比如商品标题、名称类内容),可以删掉词性校验的逻辑
- 如果需要过滤特定无意义内容,可以自行加规则匹配,比如包含特定乱码字符的直接丢弃
- 如果处理中文文本,把spaCy模型换成中文预训练模型即可,规则逻辑通用
内容的提问来源于stack exchange,提问作者solo55
相关产品推荐
相关产品推荐

