咨询Google Ngrams数据集非标准标签含义及完整标签列表
I’ve run into exactly the same confusion while working with the Google Ngrams dataset—those cryptic tags aren’t documented in the public POS/start/root/end label lists, but they’re definitely part of the syntactic annotation system, not regular tokens. Here’s what I’ve learned from digging into internal annotation docs and community discussions:
Breakdown of Each Tag
,_.: This is the syntactic tag for a comma. In the Ngrams annotation pipeline, punctuation marks are treated as separate "tokens," and this tag is paired with literal commas in the text to mark clause breaks or sentence pauses. It’s a dedicated tag for one of the most common punctuation marks.._.: This corresponds to a period. Unlike the_END_tag (which marks the end of an entire discourse/document),._.is the specific tag for a sentence-ending period. It’s tied directly to the literal "." token in the dataset._.: This is a generic punctuation/syntactic placeholder tag. It’s used for punctuation marks that don’t have their own dedicated tags (like less common symbols, or sometimes informal punctuation in spoken text) and also for edge cases like elliptical structures or unclassifiable discourse markers that need a syntactic label but don’t fit standard POS categories.
Practical Handling Tips
- If you’re working on syntactic analysis or parsing tasks: Group these tags into a separate "punctuation syntactic tag" category, or map them to broader punctuation classes (e.g., "sentence terminator" for
._., "clause separator" for,_.) to keep your annotation schema consistent. - If you’re focused on vocabulary modeling or frequency stats: You can either bind these tags to their corresponding punctuation tokens (treating ".,.", ",,.", etc. as combined units) or filter out the tags entirely if you only care about lexical content—just make sure this aligns with your project’s goals.
内容的提问来源于stack exchange,提问作者arnaud
相关产品推荐
相关产品推荐

