如何在无Doc上下文的情况下对Spacy分词结果进行反分词?
Great question! Yes, spaCy has built-in tools to reverse its tokenization logic and convert your model's output token stream back into natural, properly formatted text. This handles all the edge cases spaCy’s tokenizer accounts for (like contractions, punctuation spacing, etc.) automatically.
Here’s how to do it:
The key is to reconstruct a spaCy Doc object using your output tokens and the same spaCy model vocab you used for training tokenization. This ensures the inverse rules match exactly what was applied during initial tokenization.
import spacy # Load the SAME spaCy model you used for initial tokenization (critical!) nlp = spacy.load("en_core_web_sm") # Replace with your specific model if different # Your seq2seq model's output token list output_tokens = ["This", "does", "n't", "work", "."] # Reconstruct a spaCy Doc using the model's vocab and your tokens doc = spacy.tokens.Doc(nlp.vocab, words=output_tokens) # Get the fully reconstructed natural text reconstructed_text = doc.text print(reconstructed_text) # Output: "This doesn't work."
Why this works:
- spaCy’s
Docobject doesn’t just store tokens—it also encodes the rules for how tokens should be joined back into text (like merging "does" + "n't" into "doesn't" or removing unnecessary spaces before punctuation). - Using the same model vocab ensures you’re applying the exact inverse of the tokenization rules used during training, so the output will match natural language conventions perfectly.
Note for special tokens:
If your seq2seq model uses special tokens (like <START>, <END>, or placeholder pronouns like -PRON-), make sure to filter those out before creating the Doc object. For example:
filtered_tokens = [tok for tok in output_tokens if tok not in ["<START>", "<END>", "-PRON-"]] doc = spacy.tokens.Doc(nlp.vocab, words=filtered_tokens)
内容的提问来源于stack exchange,提问作者Sai Prasanna

