You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用purrr提取textmodel_nb输出的类别概率并转为data_frame?

Solution to Reshape textmodel_nb's PcGw Matrix into a Tidy Data Frame

Let's break down how to convert that PcGw matrix into the tidy data frame you need, using purrr and other tidyverse tools (since purrr works seamlessly with them):

Step-by-Step Code

First, make sure you have the tidyverse package loaded alongside quanteda:

library(quanteda)
library(tidyverse)

Then run your original example code to get the nb_test object, then use this code to reshape the PcGw matrix:

# Reshape the PcGw matrix into the desired data frame
tidy_probabilities <- nb_test$PcGw %>%
  # Convert matrix to data frame, keeping row names (your classes)
  as.data.frame() %>%
  # Extract row names into a dedicated "class" column
  rownames_to_column(var = "class") %>%
  # Convert from wide to long format: one row per class-feature pair
  pivot_longer(cols = -class, names_to = "variable", values_to = "probability") %>%
  # Convert back to wide format: one column per class (with P_ prefix)
  pivot_wider(names_from = class, values_from = probability, names_prefix = "P_")

What This Does

Let's walk through each step with your example:

  1. as.data.frame(): Converts the PcGw matrix (rows = classes, columns = features) into a data frame where row names are your class labels (Y/N in your example, TRUE/FALSE in your actual data).
  2. rownames_to_column(): Moves those class labels from row names into a proper column called class—this makes it easy to reshape later.
  3. pivot_longer(): Flattens the data so each row represents a single class-feature-probability combination (e.g., one row for Y + Chinese + its probability, another for N + Chinese + its probability).
  4. pivot_wider(): Pivots the data back to wide format, creating separate columns for each class (prefixed with P_ to match your P_TRUE/P_FALSE naming).

If You Want a More purrr-Focused Approach

If you prefer to lean more heavily on purrr instead of tidyr's pivot functions, you can use imap_dfr to iterate over the matrix columns directly:

tidy_probabilities <- imap_dfr(as.data.frame(nb_test$PcGw), 
                               ~ tibble(variable = .y, 
                                        class = rownames(nb_test$PcGw), 
                                        probability = .x)) %>%
  pivot_wider(names_from = class, values_from = probability, names_prefix = "P_")

Here, imap_dfr loops through each column (feature) of the matrix:

  • .y gives the column name (your feature/variable)
  • .x gives the values in the column (probabilities for each class)
    We combine these into a tibble, then pivot to wide format as before.

Result for Your Example

Running either code on your sample data will produce a data frame like this:

# A tibble: 6 × 3
  variable   P_Y      P_N     
  <chr>      <dbl>    <dbl>
1 Chinese    0.768    0.232
2 Beijing    0.113    0     
3 Shanghai   0.113    0     
4 Macao      0.113    0     
5 Tokyo      0.0377   0.384
6 Japan      0.0377   0.384

In your actual data with TRUE/FALSE classes, the columns will automatically be P_TRUE and P_FALSE instead of P_Y/P_N.

内容的提问来源于stack exchange,提问作者ℕʘʘḆḽḘ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 04:06:12