求助:基于NLTK或最优库开发疑问句识别Python脚本
哈哈,太懂这种感受了!正则对付简单问句还行,但碰到带从句的、倒装的或者口语化的问句,立马就歇菜了。下面给你几个用Python NLP工具实现的方案,从轻量到进阶都有,按需选就行:
方案1:用NLTK快速实现(轻量易上手)
NLTK是老牌NLP库,用它的词性标注就能覆盖大部分常见问句场景。先装包加下载语料:
pip install nltk
然后在代码里下载必要的模型(第一次运行需要):
import nltk from nltk.tokenize import word_tokenize from nltk.tag import pos_tag # 下载分词和词性标注的语料 nltk.download('punkt') nltk.download('averaged_perceptron_tagger') def is_question_nltk(sentence): tokens = word_tokenize(sentence) tagged_words = pos_tag(tokens) # 疑问词对应的词性标签(WDT=限定性疑问词,WP=疑问代词等) question_word_tags = {'WDT', 'WP', 'WP$', 'WRB'} # 助动词/系动词的词性标签(用来判断倒装问句,比如"Are you okay?") aux_verb_tags = {'MD', 'VB', 'VBD', 'VBG', 'VBN', 'VBP', 'VBZ'} # 情况1:句子以疑问词开头 if tagged_words[0][1] in question_word_tags: return True # 情况2:开头是助动词/系动词,后面跟着主语(倒装结构) if len(tagged_words) > 1 and tagged_words[0][1] in aux_verb_tags and tagged_words[1][1] in {'PRP', 'NN'}: return True # 情况3:口语化的反义疑问句或者带问号的句子 if tokens[-1] in {'?', 'right?', 'eh?'} or sentence.strip().endswith('?'): return True return False # 测试一波 test_cases = [ "What's the best NLP library for question detection?", "Are you planning to use this script in production?", "You love Python, don't you?", "I'm working on a text classification project.", "Could you explain how spaCy's syntax analysis works?" ] for case in test_cases: print(f"'{case}' → 是否为疑问句: {is_question_nltk(case)}")
这个方案优点是轻量、速度快,适合简单场景;缺点是对嵌套问句(比如"Do you know where I can find the docs?")这种复杂结构识别不够准。
方案2:用spaCy处理复杂句式(平衡速度与准确率)
spaCy的句法分析能力比NLTK强很多,能更好捕捉句子的深层结构,对付嵌套问句、倒装句都不在话下。先装包加模型:
pip install spacy python -m spacy download en_core_web_sm
然后写代码:
import spacy # 加载英文小模型 nlp = spacy.load("en_core_web_sm") def is_question_spacy(sentence): doc = nlp(sentence) # 先检查最基础的:末尾带问号 if doc.text.strip().endswith('?'): return True # 遍历句法依赖关系,找疑问词或者倒装结构 for token in doc: # 疑问词的依赖标签或者词性标签 if token.tag_ in {'WDT', 'WP', 'WP$', 'WRB'}: return True # 如果根节点是助动词,且主语在它后面(倒装结构,比如"Can you help me?") if token.dep_ == 'ROOT' and token.pos_ == 'AUX': subj = [child for child in token.children if child.dep_ == 'nsubj'] if subj and token.i < subj[0].i: return True # 处理反义疑问句 if any(token.text.lower() in {'don', 'didn', 'aren', 'isn'} and token.dep_ == 'aux' for token in doc): return True return False # 加个复杂测试用例 test_cases.append("Do you know where the nearest coffee shop is?") for case in test_cases: print(f"'{case}' → 是否为疑问句: {is_question_spacy(case)}")
这个方案能搞定大部分复杂场景,而且速度也不错,是我个人比较推荐的折中方案。
方案3:用预训练模型(极致准确率)
如果对准确率要求极高,比如要处理口语化文本、隐晦的反问句,那直接上Hugging Face的预训练模型就行,不需要自己训练数据,用zero-shot分类就能搞定:
pip install transformers
代码示例:
from transformers import pipeline # 加载Bart的预训练模型,支持zero-shot分类 classifier = pipeline("zero-shot-classification", model="facebook/bart-large-mnli") def is_question_transformers(sentence): # 定义分类标签:疑问句和陈述句 candidate_labels = ["question", "statement"] result = classifier(sentence, candidate_labels) # 返回得分最高的标签是否为question return result['labels'][0] == "question" # 测试所有用例 for case in test_cases: print(f"'{case}' → 是否为疑问句: {is_question_transformers(case)}")
这个方案准确率拉满,但缺点是模型大,运行速度慢,适合离线处理或者对精度要求极高的场景。
选择建议
- 简单场景、追求速度:选NLTK
- 中等复杂度、平衡速度与准确率:选spaCy
- 极致准确率、复杂文本:选Transformers预训练模型
内容的提问来源于stack exchange,提问作者Freakant
相关产品推荐
相关产品推荐

