You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在无Doc上下文的情况下对Spacy分词结果进行反分词?

Reverse spaCy Tokenization for Seq2Seq Outputs

Great question! Yes, spaCy has built-in tools to reverse its tokenization logic and convert your model's output token stream back into natural, properly formatted text. This handles all the edge cases spaCy’s tokenizer accounts for (like contractions, punctuation spacing, etc.) automatically.

Here’s how to do it:

The key is to reconstruct a spaCy Doc object using your output tokens and the same spaCy model vocab you used for training tokenization. This ensures the inverse rules match exactly what was applied during initial tokenization.

import spacy

# Load the SAME spaCy model you used for initial tokenization (critical!)
nlp = spacy.load("en_core_web_sm")  # Replace with your specific model if different

# Your seq2seq model's output token list
output_tokens = ["This", "does", "n't", "work", "."]

# Reconstruct a spaCy Doc using the model's vocab and your tokens
doc = spacy.tokens.Doc(nlp.vocab, words=output_tokens)

# Get the fully reconstructed natural text
reconstructed_text = doc.text
print(reconstructed_text)  # Output: "This doesn't work."

Why this works:

  • spaCy’s Doc object doesn’t just store tokens—it also encodes the rules for how tokens should be joined back into text (like merging "does" + "n't" into "doesn't" or removing unnecessary spaces before punctuation).
  • Using the same model vocab ensures you’re applying the exact inverse of the tokenization rules used during training, so the output will match natural language conventions perfectly.

Note for special tokens:

If your seq2seq model uses special tokens (like <START>, <END>, or placeholder pronouns like -PRON-), make sure to filter those out before creating the Doc object. For example:

filtered_tokens = [tok for tok in output_tokens if tok not in ["<START>", "<END>", "-PRON-"]]
doc = spacy.tokens.Doc(nlp.vocab, words=filtered_tokens)

内容的提问来源于stack exchange,提问作者Sai Prasanna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:35:47