Mallet主题建模:移除高频与低频词的操作步骤问询
Hey there! Great question—pruning low and high frequency terms is absolutely a smart strategy to refine your topic models, especially in art history where terminology can range from hyper-niche to overly broad. Your hunch about using Mallet's prune commands and prune-document-freq parameter is totally correct, and I’ll walk you through the full, command-line-only workflow (no Java required!) plus some extra tips tailored to your field.
First: Confirm Your Core Tools Are Correct
Yes, vectors2vectors with --prune-count (for low-frequency terms) and --prune-document-freq (for high-frequency terms) is exactly the way to follow D. Mimno’s advice. These tools let you clean up your corpus after initial import, which is far more flexible than filtering during the import step.
Full Step-by-Step Workflow
Assuming you already have your art history corpus (individual text files, or a single file with one document per line), here’s the complete process:
1. Import Your Raw Corpus into Mallet Format
First, convert your text files into Mallet’s internal .mallet format (if you haven’t already). This step also lets you apply your existing stopword list upfront:
bin/mallet import-dir \ --input /path/to/your/art-history-corpus \ --output original_corpus.mallet \ --keep-sequence \ --remove-stopwords \ --stoplist-file /path/to/your/custom-stopwords.txt
--keep-sequence: Preserves word order (required for proper topic modeling).--remove-stopwords: Uses your existing stopword list to filter common, non-informative words first.--stoplist-file: Points to your pre-made stopword text file (one word per line).
2. Prune Low and High Frequency Terms with vectors2vectors
This is the key step to implement Mimno’s recommendation. Run this command to create a cleaned, pruned corpus:
bin/mallet vectors2vectors \ --input original_corpus.mallet \ --output pruned_corpus.mallet \ --prune-count 10 \ --prune-document-freq 0.9 1.0
Let’s break down the critical parameters:
--prune-count 10: Removes any word that appears 10 times or fewer across the entire corpus. Adjust this number based on your corpus size—use 20 for a very large dataset, or 5 if your corpus is small and niche.--prune-document-freq 0.9 1.0: Removes words that appear in 90% to 100% of documents (these are overly broad terms like "art" or "century" that don’t add topic specificity). If you prefer absolute document counts instead of percentages, replace the range with numbers (e.g.,500 1000000to remove words appearing in 500+ documents).
3. Train Your Topic Model with the Pruned Corpus
Now use the cleaned corpus to train your model—this should produce more focused, interpretable topics aligned with your art history research:
bin/mallet train-topics \ --input pruned_corpus.mallet \ --num-topics 25 \ # Adjust to your desired number of topics (e.g., 20-30 for medium corpora) --output-state topic-state.gz \ --output-topic-keys topic-keys.txt \ --output-doc-topics doc-topics.txt \ --num-iterations 1000 # Increase to 1500 if your model isn't converging
- The output files (
topic-keys.txt,doc-topics.txt) will let you analyze your topics and document-topic assignments.
Alternative Methods to Consider
While vectors2vectors is the most straightforward approach, you have a couple of other options:
- Manual High-Frequency Stopwords: Generate a list of the most common words in your corpus (using the
print-vectorscommand below) and add domain-specific broad terms to your stopword list, then re-import the corpus. This gives you full control to keep critical high-frequency terms (e.g., "renaissance" even if it’s common). - Filter During Import: Use
--token-regexin theimport-dircommand to exclude very short words or non-art-history terms, but this is less precise than post-import pruning.
Pro Tips for Art History Research
- Check Word Frequencies First: Before pruning, run this command to get a full list of word counts and document frequencies—this helps you set smarter thresholds and avoid pruning critical domain terms:
Openbin/mallet print-vectors --input original_corpus.mallet --output word-frequencies.txtword-frequencies.txtto spot terms like "fresco" or "iconography" that might be mid-frequency but central to your research—don’t accidentally cut these! - Iterate and Compare: Train models with different pruning thresholds (e.g.,
--prune-count 5vs15) and compare the topic keys to see which produces the most meaningful results for your project.
内容的提问来源于stack exchange,提问作者Eugenie

