You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

SQuAD QA答案含BERT特殊字符的含义及评估处理方式咨询

Handling Invalid QA Predictions with [CLS]/[SEP] Tokens in SQuAD v2.0

Great question—this is a super common pain point when working with SQuAD v2.0 and transformer-based QA models, especially since the task explicitly expects you to handle both answerable and unanswerable questions correctly. Let’s break this down clearly:

First, Recap SQuAD v2.0’s Core Rule

SQuAD v2.0 allows questions that have no valid answer in the provided context. For these, the model is expected to output an empty string to be scored correctly by the official evaluation script.

How to Handle Your Specific Case

Your example where the prediction includes [CLS] and [SEP] is not a valid answer—those special tokens are part of the transformer’s input formatting, not part of the original context text. Here’s how to approach this:

1. Distinguish Between Explicit Unanswerable Predictions and Invalid Spans

  • When start/end scores both point to [CLS] (index 0): This is the model’s explicit signal that it thinks the question is unanswerable. You’re already handling this correctly by converting to an empty string.
  • When predictions include [CLS]/[SEP] or fall outside the context token range: This is an invalid prediction. The model failed to locate a valid span within the actual context text, so you should treat this the same as an unanswerable question—convert it to an empty string before running the evaluation script.

2. Why This Matters for Evaluation

The official SQuAD v2.0 evaluation script (evaluate-v2.0.py) compares your model’s output directly against ground-truth answers. If you submit predictions with [CLS]/[SEP], the script will treat those as literal answer strings. Since the ground-truth will never include these tokens, all such predictions will be marked as incorrect, which artificially drags down your model’s performance metrics (exact match, F1 score). This doesn’t reflect your model’s actual ability to find valid answers—it just penalizes formatting errors.

3. Updated Code to Fix This

Here’s how to modify your existing code to filter out invalid spans:

tokenizer = AutoTokenizer.from_pretrained("ktrapeznikov/albert-xlarge-v2-squad-v2")
model = AutoModelForQuestionAnswering.from_pretrained("ktrapeznikov/albert-xlarge-v2-squad-v2")
question = "Why aren't the examples of bouregois architecture visible today?"
text = """Exceptional examples of the bourgeois architecture of the later periods were not restored by the communist authorities after the war (like mentioned Kronenberg Palace and Insurance Company Rosja building) or they were rebuilt in socialist realism style (like Warsaw Philharmony edifice originally inspired by Palais Garnier in Paris). Despite that the Warsaw University of Technology building (1899–1902) is the most interesting of the late 19th-century architecture. Some 19th-century buildings in the Praga district (the Vistula’s right bank) have been restored although many have been poorly maintained. Warsaw’s municipal government authorities have decided to rebuild the Saxon Palace and the Brühl Palace, the most distinctive buildings in prewar Warsaw."""

input_dict = tokenizer.encode_plus(question, text, return_tensors="pt")
input_ids = input_dict["input_ids"].tolist()
start_scores, end_scores = model(**input_dict)
all_tokens = tokenizer.convert_ids_to_tokens(input_ids[0])

# Locate the bounds of the actual context text in the token list
sep_indices = [i for i, token in enumerate(all_tokens) if token == '[SEP]']
context_start_idx = sep_indices[0] + 1  # After the first [SEP] (end of question)
context_end_idx = sep_indices[1] - 1    # Before the final [SEP] (end of context)

start_pred = torch.argmax(start_scores).item()
end_pred = torch.argmax(end_scores).item()

# Check if the predicted span is within the valid context range
if context_start_idx <= start_pred <= end_pred <= context_end_idx:
    # Valid span: clean up and format the answer
    answer = ' '.join(all_tokens[start_pred : end_pred + 1]).replace('▁', '')
else:
    # Invalid span: treat as unanswerable
    answer = ''

print(answer)

4. Extra Context on Model Behavior

Sometimes models output invalid spans like this if they weren’t fine-tuned perfectly, or if the input context has edge cases. Adding this post-processing step ensures you’re evaluating the model’s actual QA ability, not its tendency to produce formatting artifacts.

内容的提问来源于stack exchange,提问作者Pedram

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 15:39:08