关于PTBTokenizer识别无空格句末标点为独立分词的设置咨询
Great question! I’ve dealt with exactly this issue when using PTBTokenizer on messy social media text like Twitter—non-standard formatting (like missing spaces after sentence-ending punctuation) throws off the tokenizer’s default rules.
Unfortunately, PTBTokenizer (specifically the common NLTK implementation based on the Penn Treebank rules) doesn’t have a built-in configuration switch to handle this edge case directly. But there are two straightforward workarounds that’ll fix the problem:
1. Preprocess Your Text to Insert Missing Spaces
The quickest fix is to run a regex-based preprocessing step on your Twitter text before feeding it to the tokenizer. This will add a space between sentence-ending punctuation (., !, ?) and a capitalized word that follows it without a space.
For example, in Python, you can use this regex substitution:
import re def preprocess_text(text): # Match lowercase letter + punctuation + uppercase letter, add space return re.sub(r'([a-z])([.!?])([A-Z])', r'\1\2 \3', text) # Test it out raw_text = "This is a test.And some more." processed_text = preprocess_text(raw_text) # Output: "This is a test. And some more."
Once you run this preprocessing, PTBTokenizer will correctly split test. and And into separate tokens: ["This", "is", "a", "test", ".", "And", "some", "more", "."].
You can tweak the regex to cover other edge cases too—like if you see numbers followed by punctuation and a capitalized word, adjust the pattern to include [0-9].
2. Customize PTBTokenizer's Splitting Rules
If you want a more integrated solution, you can modify the underlying rules of the TreebankWordTokenizer (which powers NLTK’s PTBTokenizer). The tokenizer uses a list of regex patterns to split text; you can add a new pattern to split tokens like test.And into three parts.
Here’s how you’d adjust it in NLTK:
from nltk.tokenize import TreebankWordTokenizer # Create a custom tokenizer with added split rule custom_tokenizer = TreebankWordTokenizer() # Add a new regex pattern to split word.punctuationCapitalWord custom_tokenizer.PUNCTUATION += [ (r'(\w+)([.!?])([A-Z]\w+)', r'\1 \2 \3') ] # Test the custom tokenizer raw_text = "This is a test.And some more." tokens = custom_tokenizer.tokenize(raw_text) # Output: ["This", "is", "a", "test", ".", "And", "some", "more", "."]
This approach embeds the fix directly into the tokenizer, so you don’t have to run a separate preprocessing step every time.
Final Notes
Twitter text is full of these non-standard formatting quirks, so combining preprocessing with a slightly tweaked tokenizer usually gives the best results. For most cases, the preprocessing step is simpler and easier to maintain.
内容的提问来源于stack exchange,提问作者Berfin Aktaş

