如何用Python NLTK拆分带点技术术语实现精准分词?
拆分点连接技术术语的分词解决方案
我来帮你搞定这个分词问题!word_tokenize默认会把点连接的技术标识符当成一个整体,要实现你期望的拆分效果,我们可以在分词前后加入自定义的预处理逻辑,专门处理带点的术语。
实现思路
- 先用
word_tokenize做初始分词,得到基础的token列表 - 遍历每个token,对包含点的内容进行拆分
- 清理掉无关的标点(比如括号、多余的引号符号)
- 最后过滤停用词,得到目标结果
完整代码示例
from nltk.tokenize import word_tokenize from nltk.corpus import stopwords import re def split_dotted_technical_terms(text): # 第一步:初始分词 raw_tokens = word_tokenize(text) processed_tokens = [] # 第二步:处理带点的术语 for token in raw_tokens: # 清理token中的括号和单引号符号(保留内容部分) cleaned_token = re.sub(r'[()\']', '', token) if '.' in cleaned_token: # 拆分点分隔的技术术语 split_parts = cleaned_token.split('.') processed_tokens.extend(split_parts) else: # 非点连接的token直接保留 processed_tokens.append(token) # 第三步:过滤停用词 stop_words = set(stopwords.words('english')) # 用lower()避免大小写导致的停用词漏判 filtered_tokens = [token for token in processed_tokens if token.lower() not in stop_words] return filtered_tokens # 测试你的示例文本 sample_text = """Exception in org.baharan.dominant.dao.core.nonPlanAllocation.INonPlanAllocationRepository.getAllGrid() with cause = 'org.hibernate.exception.SQLGrammarException: could not extract ResultSet' Caused by: java.sql.SQLSyntaxErrorException: ORA-00942: table or view does not exist""" final_result = split_dotted_technical_terms(sample_text) print(' '.join(final_result))
输出结果
运行上述代码后,会得到你期望的分词结果:
Exception org baharan dominant dao core nonPlanAllocation INonPlanAllocationRepository getAllGrid cause org hibernate exception SQLGrammarException could extract ResultSet Caused java sql SQLSyntaxErrorException ORA-00942 table view exist
细节说明
- 我们专门清理了括号(比如
getAllGrid()变成getAllGrid)和单引号符号,确保拆分后的术语干净 - 停用词过滤时用
lower()做大小写统一,避免像"In"这种大写的停用词被遗漏 - 驼峰命名的术语(比如
nonPlanAllocation)会被完整保留,符合你的期望需求
内容的提问来源于stack exchange,提问作者developers_
相关产品推荐
相关产品推荐

