You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用TFIDFVectorizer处理Cranfield数据集出现拼接词问题求解

问题描述

我正在使用Cranfield数据集构建索引器与查询处理器,采用TFIDFVectorizer对数据进行分词。但使用后查看词汇表时,发现大量由两个单词拼接而成的token,比如freevibration、slendersharp这类。

我使用的代码如下:

import re
from sklearn.feature_extraction import text
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
import numpy as np
from nltk import word_tokenize          
from nltk.stem import WordNetLemmatizer
#reading the data
with open('cran.all', 'r') as f:
    content_string=""
    content = [line.replace('\n','') for line in f]
    content =  content_string.join(content)
    doc=re.split('.I\s[0-9]{1,4}',content)
    f.close()
#some data cleaning
doc = [line.replace('.T',' ').replace('.B',' ').replace('.A',' ').replace('.W',' ') for line in doc]
del doc[0]
doc= [ re.sub('[^A-Za-z]+', ' ', lines) for lines in doc]

vectorizer = TfidfVectorizer(analyzer ='word', ngram_range=(1,1), stop_words=text.ENGLISH_STOP_WORDS,lowercase=True)
X = vectorizer.fit_transform(doc)
print(vectorizer.vocabulary_)

词汇表部分示例:

'freevibration': 7222, 'slendersharp': 15197, 'frequentlyapproximated': 7249, 'notapplicable': 11347, 'rateof': 13727, 'itsvalue': 9443, 'speedflow': 15516, 'movingwith': 11001, 'speedsolution': 15531, 'centerof': 3314, 'hypersoniclow': 8230, 'neice': 11145, 'rutkowski': 14444, 'chann': 3381, 'layerapproximations': 9828, 'probsteinhave': 13353, 'thishypersonic': 17752

处理小数据时未出现该问题,请问如何避免这种情况?

解决方案

问题根源出在数据读取和预处理阶段,具体修正方式如下:

  1. 修复换行符处理逻辑
    原代码把换行符直接替换为空字符串,导致分属两行的单词被硬拼接。需将换行符替换为空格:

    content = [line.replace('\n',' ') for line in f]
    
  2. 修正正则表达式的转义问题
    正则中的.是通配符,不是字面点号,会导致文档分割错误。给.添加转义符:

    doc=re.split('\.I\s[0-9]{1,4}',content)
    
  3. 可选优化:使用自定义分词器
    结合NLTK的分词和词形还原工具,进一步避免异常token:

    class LemmaTokenizer:
        def __init__(self):
            self.lemmatizer = WordNetLemmatizer()
        def __call__(self, doc):
            return [self.lemmatizer.lemmatize(t) for t in word_tokenize(doc) if t.isalpha()]
    
    # 替换原vectorizer初始化代码
    vectorizer = TfidfVectorizer(tokenizer=LemmaTokenizer(), stop_words=text.ENGLISH_STOP_WORDS, lowercase=True)
    

内容的提问来源于stack exchange,提问作者SHIVANSHU SAHOO

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 09:02:57