You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:基于NLTK或最优库开发疑问句识别Python脚本

哈哈,太懂这种感受了!正则对付简单问句还行,但碰到带从句的、倒装的或者口语化的问句,立马就歇菜了。下面给你几个用Python NLP工具实现的方案,从轻量到进阶都有,按需选就行:

方案1:用NLTK快速实现(轻量易上手)

NLTK是老牌NLP库,用它的词性标注就能覆盖大部分常见问句场景。先装包加下载语料:

pip install nltk

然后在代码里下载必要的模型(第一次运行需要):

import nltk
from nltk.tokenize import word_tokenize
from nltk.tag import pos_tag

# 下载分词和词性标注的语料
nltk.download('punkt')
nltk.download('averaged_perceptron_tagger')

def is_question_nltk(sentence):
    tokens = word_tokenize(sentence)
    tagged_words = pos_tag(tokens)
    
    # 疑问词对应的词性标签(WDT=限定性疑问词,WP=疑问代词等)
    question_word_tags = {'WDT', 'WP', 'WP$', 'WRB'}
    # 助动词/系动词的词性标签(用来判断倒装问句,比如"Are you okay?")
    aux_verb_tags = {'MD', 'VB', 'VBD', 'VBG', 'VBN', 'VBP', 'VBZ'}
    
    # 情况1:句子以疑问词开头
    if tagged_words[0][1] in question_word_tags:
        return True
    # 情况2:开头是助动词/系动词,后面跟着主语(倒装结构)
    if len(tagged_words) > 1 and tagged_words[0][1] in aux_verb_tags and tagged_words[1][1] in {'PRP', 'NN'}:
        return True
    # 情况3:口语化的反义疑问句或者带问号的句子
    if tokens[-1] in {'?', 'right?', 'eh?'} or sentence.strip().endswith('?'):
        return True
    return False

# 测试一波
test_cases = [
    "What's the best NLP library for question detection?",
    "Are you planning to use this script in production?",
    "You love Python, don't you?",
    "I'm working on a text classification project.",
    "Could you explain how spaCy's syntax analysis works?"
]

for case in test_cases:
    print(f"'{case}' → 是否为疑问句: {is_question_nltk(case)}")

这个方案优点是轻量、速度快,适合简单场景;缺点是对嵌套问句(比如"Do you know where I can find the docs?")这种复杂结构识别不够准。

方案2:用spaCy处理复杂句式(平衡速度与准确率)

spaCy的句法分析能力比NLTK强很多,能更好捕捉句子的深层结构,对付嵌套问句、倒装句都不在话下。先装包加模型:

pip install spacy
python -m spacy download en_core_web_sm

然后写代码:

import spacy

# 加载英文小模型
nlp = spacy.load("en_core_web_sm")

def is_question_spacy(sentence):
    doc = nlp(sentence)
    # 先检查最基础的:末尾带问号
    if doc.text.strip().endswith('?'):
        return True
    # 遍历句法依赖关系,找疑问词或者倒装结构
    for token in doc:
        # 疑问词的依赖标签或者词性标签
        if token.tag_ in {'WDT', 'WP', 'WP$', 'WRB'}:
            return True
        # 如果根节点是助动词,且主语在它后面(倒装结构,比如"Can you help me?")
        if token.dep_ == 'ROOT' and token.pos_ == 'AUX':
            subj = [child for child in token.children if child.dep_ == 'nsubj']
            if subj and token.i < subj[0].i:
                return True
    # 处理反义疑问句
    if any(token.text.lower() in {'don', 'didn', 'aren', 'isn'} and token.dep_ == 'aux' for token in doc):
        return True
    return False

# 加个复杂测试用例
test_cases.append("Do you know where the nearest coffee shop is?")

for case in test_cases:
    print(f"'{case}' → 是否为疑问句: {is_question_spacy(case)}")

这个方案能搞定大部分复杂场景,而且速度也不错,是我个人比较推荐的折中方案。

方案3:用预训练模型(极致准确率)

如果对准确率要求极高,比如要处理口语化文本、隐晦的反问句,那直接上Hugging Face的预训练模型就行,不需要自己训练数据,用zero-shot分类就能搞定:

pip install transformers

代码示例:

from transformers import pipeline

# 加载Bart的预训练模型,支持zero-shot分类
classifier = pipeline("zero-shot-classification", model="facebook/bart-large-mnli")

def is_question_transformers(sentence):
    # 定义分类标签:疑问句和陈述句
    candidate_labels = ["question", "statement"]
    result = classifier(sentence, candidate_labels)
    # 返回得分最高的标签是否为question
    return result['labels'][0] == "question"

# 测试所有用例
for case in test_cases:
    print(f"'{case}' → 是否为疑问句: {is_question_transformers(case)}")

这个方案准确率拉满,但缺点是模型大,运行速度慢,适合离线处理或者对精度要求极高的场景。

选择建议
  • 简单场景、追求速度:选NLTK
  • 中等复杂度、平衡速度与准确率:选spaCy
  • 极致准确率、复杂文本:选Transformers预训练模型

内容的提问来源于stack exchange,提问作者Freakant

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:13:38