quanteda中dfm_groups无分组作用及dfm_weight(prop)频率含义问询
Let’s break down your two questions clearly:
1. Is dfm_group(groups=...) with no grouping variable equivalent to removing that line?
When you call plot_topterms without passing a grouping variable, dfm_group(groups=...) will default to grouping all documents into a single group named all.
In terms of frequency calculation results, this is almost identical to removing the dfm_group line entirely—both approaches treat the entire corpus as one unified set when calculating proportional weights. The only minor difference is structural:
- If you remove
dfm_group, the output fromtextstat_frequencywon’t have agroupcolumn. - If you keep
dfm_groupwithout a grouping variable, you’ll get agroupcolumn with every row set toall, which is why yourfacet_wraprenders a single panel labeled "all".
For your use case (plotting top terms across a single "all" group), the end visual and numerical results will be the same either way.
2. Does a relative frequency of 60 for "fun" mean 60% of documents contain the word?
No, that’s a common misconception! Let’s clarify what dfm_weight(scheme='prop') actually calculates:
scheme='prop'computes the term’s relative frequency within the entire corpus (or group)—meaning it’s the number of times the term appears divided by the total number of term occurrences in that corpus/group, multiplied by 100 (which is why you see 60 instead of 0.6).- This is very different from document inclusion rate (the percentage of documents that contain the term at least once). To get that metric, you’d need to use
dfm_weight(scheme='docfreq')instead, which counts how many documents each term appears in, then normalize by total documents.
For example:
- If your corpus has 100 total term occurrences, and "fun" appears 60 times, its proportional weight will be 60.
- But if only 30 out of 100 documents contain "fun" (even if it’s repeated multiple times in those 30), the document inclusion rate would be 30%, not 60%.
内容的提问来源于stack exchange,提问作者Ted Mosby

