如何用purrr提取textmodel_nb输出的类别概率并转为data_frame?
textmodel_nb's PcGw Matrix into a Tidy Data Frame Let's break down how to convert that PcGw matrix into the tidy data frame you need, using purrr and other tidyverse tools (since purrr works seamlessly with them):
Step-by-Step Code
First, make sure you have the tidyverse package loaded alongside quanteda:
library(quanteda) library(tidyverse)
Then run your original example code to get the nb_test object, then use this code to reshape the PcGw matrix:
# Reshape the PcGw matrix into the desired data frame tidy_probabilities <- nb_test$PcGw %>% # Convert matrix to data frame, keeping row names (your classes) as.data.frame() %>% # Extract row names into a dedicated "class" column rownames_to_column(var = "class") %>% # Convert from wide to long format: one row per class-feature pair pivot_longer(cols = -class, names_to = "variable", values_to = "probability") %>% # Convert back to wide format: one column per class (with P_ prefix) pivot_wider(names_from = class, values_from = probability, names_prefix = "P_")
What This Does
Let's walk through each step with your example:
as.data.frame(): Converts thePcGwmatrix (rows = classes, columns = features) into a data frame where row names are your class labels (Y/Nin your example,TRUE/FALSEin your actual data).rownames_to_column(): Moves those class labels from row names into a proper column calledclass—this makes it easy to reshape later.pivot_longer(): Flattens the data so each row represents a single class-feature-probability combination (e.g., one row forY+Chinese+ its probability, another forN+Chinese+ its probability).pivot_wider(): Pivots the data back to wide format, creating separate columns for each class (prefixed withP_to match yourP_TRUE/P_FALSEnaming).
If You Want a More purrr-Focused Approach
If you prefer to lean more heavily on purrr instead of tidyr's pivot functions, you can use imap_dfr to iterate over the matrix columns directly:
tidy_probabilities <- imap_dfr(as.data.frame(nb_test$PcGw), ~ tibble(variable = .y, class = rownames(nb_test$PcGw), probability = .x)) %>% pivot_wider(names_from = class, values_from = probability, names_prefix = "P_")
Here, imap_dfr loops through each column (feature) of the matrix:
.ygives the column name (your feature/variable).xgives the values in the column (probabilities for each class)
We combine these into a tibble, then pivot to wide format as before.
Result for Your Example
Running either code on your sample data will produce a data frame like this:
# A tibble: 6 × 3 variable P_Y P_N <chr> <dbl> <dbl> 1 Chinese 0.768 0.232 2 Beijing 0.113 0 3 Shanghai 0.113 0 4 Macao 0.113 0 5 Tokyo 0.0377 0.384 6 Japan 0.0377 0.384
In your actual data with TRUE/FALSE classes, the columns will automatically be P_TRUE and P_FALSE instead of P_Y/P_N.
内容的提问来源于stack exchange,提问作者ℕʘʘḆḽḘ

