You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按可变样本量从DataFrame为不同说话者生成月度随机词汇表

解决按月份为不同说话者抽取指定数量词汇的问题

错误原因

你遇到的报错核心问题是:

  • list1中的每个元素是对应说话者的24个月完整数据框,element$vocab_size是长度为24的向量,但slice_sample()的n参数要求是单个常数,不能传入向量。
  • 后续出现的$ operator is invalid for atomic vectors是因为尝试对非数据框结构使用$索引,本质还是没有逐行处理每个月份的样本量。

解决方案

下面提供两种可行实现方式,分别基于tidyverse工具链和基础R:

方法一:使用tidyverse(dplyr + purrr)

这种方式逻辑清晰,结果结构规整,适合后续分析:

首先给lexicon添加列名(避免无列名的麻烦):

lexicon <- data.frame(
  word = c("a", "about", "above", "ain't", "all", "am", "an", "and",  
           "animal",  "ankle", "ant" ,"any", "apple","applesauce", 
           "asleep", "at",  "ate",  "aunt", "auntie",  "aunty's", 
           "awake", "away", "baa", "baby" , "baby+doll", "bad" ,
           "ball", "balloon", "banana", "basket", "bat", "bath", 
           "bathing", "bathtub", "be", "beach", "bead", "bean",
           "because", "bed", "beddy", "bee", "been", "behind",
           "being", "belt", "bench", "bib", "bicycle", "big")
)

然后按说话者分组嵌套,逐月份抽取词汇:

library(dplyr)
library(purrr)

# 对df1按说话者分组,嵌套月份数据
df_nested <- df1 %>%
  group_by(Speaker) %>%
  nest()

# 逐说话者、逐月份抽取对应数量的词汇
vocab_data <- df_nested %>%
  mutate(
    monthly_vocab = map(data, function(month_data) {
      # 对每个月份的样本量,抽取词汇(处理样本量为0的情况)
      map(month_data$vocab_size, function(n) {
        if (n == 0) {
          tibble(word = character(0))
        } else {
          slice_sample(lexicon, n = n, replace = TRUE)
        }
      }) %>%
        # 绑定月份信息与对应词汇
        bind_cols(month_data %>% select(months), .) %>%
        rename(monthly_words = ...2)
    })
  ) %>%
  unnest(monthly_vocab)

最终vocab_data是一个数据框,包含Speaker、months、monthly_words三列,其中monthly_words是每个月份抽取的词汇列表(若要展开为每行一个词汇,可再加一层unnest(monthly_words))。

方法二:使用基础R

如果偏好基础R语法,可通过嵌套循环实现:

# 先给lexicon加列名
lexicon <- data.frame(
  word = c("a", "about", "above", "ain't", "all", "am", "an", "and",  
           "animal",  "ankle", "ant" ,"any", "apple","applesauce", 
           "asleep", "at",  "ate",  "aunt", "auntie",  "aunty's", 
           "awake", "away", "baa", "baby" , "baby+doll", "bad" ,
           "ball", "balloon", "banana", "basket", "bat", "bath", 
           "bathing", "bathtub", "be", "beach", "bead", "bean",
           "because", "bed", "beddy", "bee", "been", "behind",
           "being", "belt", "bench", "bib", "bicycle", "big")
)

# 处理每个说话者的数据
vocab_data <- lapply(list1, function(speaker_df) {
  # 逐月份处理
  lapply(1:nrow(speaker_df), function(row_idx) {
    n_samples <- speaker_df$vocab_size[row_idx]
    # 处理样本量为0的情况
    if (n_samples == 0) {
      result <- data.frame(word = character(0))
    } else {
      # 基础R的随机抽样
      result <- lexicon[sample(nrow(lexicon), n_samples, replace = TRUE), , drop = FALSE]
    }
    # 添加月份信息
    cbind(months = speaker_df$months[row_idx], result)
  })
})

最终vocab_data是一个列表,每个元素对应一位说话者,内部包含24个数据框,每个数据框存储对应月份的词汇和月份编号。

内容的提问来源于stack exchange,提问作者Catherine Laing

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 23:43:10