为何mlr的friedmanTestBMR要求至少2个分类器?统计原因问询
Why
friedmanTestBMR() Requires ≥2 Classifiers but posthoc.friedman.nemenyi.test() Doesn't? Great question! Let's unpack this from both statistical fundamentals and package design choices:
1. Statistical Core of the Friedman-Nemenyi Test
First, let's recall what this test is meant for:
- The Friedman test is an overall non-parametric test to check if there are significant differences between ≥2 related groups (in your case, different classifiers) across multiple blocks (your tasks/data subsets). Its null hypothesis is "all classifiers perform equally well across tasks".
- The Nemenyi post-hoc test then follows up on a significant Friedman result to pinpoint exactly which pairs of classifiers differ.
If you only have one classifier, there’s nothing to compare against. The test’s entire premise relies on having multiple models to contrast—running it on a single model doesn’t yield any meaningful statistical inference. The null hypothesis becomes trivial (there’s only one group to evaluate), and the post-hoc test has no pairs to examine.
2. Package Design Differences
mlr::friedmanTestBMR()is built specifically forBenchmarkResultobjects, which are designed to store results from multiple models tested on multiple tasks. The mlr team added a hard check here because they want to enforce statistical validity upfront—they don’t want users wasting time running tests that can’t produce meaningful conclusions.PMCMR::posthoc.friedman.nemenyi.test()is a lower-level, more flexible function that focuses on numerical computation rather than statistical validity checks. It will execute if you pass it the right data structure, even if that structure violates the test’s intended use case. It’s up to the user to ensure their input makes statistical sense.
3. What to Do Next for Your Experiment
If your goal is to analyze performance differences across your tasks with a single classifier, the Friedman-Nemenyi test isn’t the right tool. Instead:
- If you want to check if your model’s performance varies significantly across tasks, use a test like the Kruskal-Wallis H-test (for independent task subsets) or a repeated-measures ANOVA (if your tasks are paired/related and data meets normality assumptions).
- If you specifically need to use Friedman-Nemenyi to compare models, you’ll need to add at least one more classifier (e.g., a different random forest implementation, XGBoost, or SVM) as a baseline to contrast against your ranger model.
内容的提问来源于stack exchange,提问作者Edvin
相关产品推荐
相关产品推荐

