咨询:去除带撇号字符串detokenize后的多余空格的方法
Hey there! I totally get how annoying it is when your detokenization step splits common contractions like can't or won't into messy can ' t or won ' t. Let's cover a few reliable ways to fix this, depending on your needs:
1. Quick Fix: Regular Expression Replacement
If you want a lightweight solution without adding extra libraries, a regex can target exactly these split contractions and stitch them back together.
The pattern looks for a word, followed by a space + apostrophe + space, then another short word (usually t for most contractions, but this works for others like don ' t too). Here's the code:
import re # Your original detokenized output with split contractions raw_output = "i can ' t use this cause they won ' t fit" # Fix the contractions fixed_output = re.sub(r"(\w+) ' (\w+)", r"\1'\2", raw_output) # Result: "i can't use this cause they won't fit"
This regex works for most common contractions, and it'll also fix split possessives like John ' s → John's which is a nice bonus. Just note it won't handle edge cases like rare contractions, but it's perfect for everyday use.
2. Robust Solution: Use a Purpose-Built Detokenizer
If you're working with more complex text (not just contractions), using a detokenizer designed for English text will handle this and other formatting issues automatically. The NLTK library's TreebankWordDetokenizer is great for this—it's built to reverse the Treebank tokenization standard, which intentionally splits contractions during tokenization.
First, install and set up NLTK:
import nltk nltk.download('punkt') # Only needed once from nltk.tokenize.treebank import TreebankWordDetokenizer # Your original token list (tk) tk = ["i", "can", "'", "t", "use", "this", "cause", "they", "won", "'", "t", "fit"] # Detokenize correctly fixed_output = TreebankWordDetokenizer().detokenize(tk) # Result: "i can't use this cause they won't fit"
This method doesn't just fix contractions—it'll also handle things like removing spaces before commas/periods (hello , world → hello, world) and other standard punctuation formatting.
3. Alternative: SpaCy Detokenization
If you're already using SpaCy for NLP tasks, you can leverage its tokenizer to rebuild the text correctly:
import spacy # Load the English model (install with pip install spacy && python -m spacy download en_core_web_sm) nlp = spacy.load("en_core_web_sm") # Convert your token list to a SpaCy Doc doc = nlp.tokenizer.tokens_from_list(tk) fixed_output = doc.text # Result: "i can't use this cause they won't fit"
SpaCy's tokenizer understands English grammar rules, so it'll automatically merge the split contraction tokens back into their correct form.
内容的提问来源于stack exchange,提问作者BenT

