You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

运行R抽样函数时遭遇sample.int无效首参数错误求助

问题描述

我运行一个抽样函数时反复遇到无法定位的错误,错误信息如下:

Error in sample.int(length(x), size, replace, prob) :
invalid first argument

我的代码如下:

idpSample <- function(x) {

main.sample <- sample(
idp.frame[COUNTY == idp.breakdown[x, County], GEOID9],
idp.breakdown[x, idp_cluster],
prob = idp.frame[COUNTY == idp.breakdown[x, County], mt12_est_idps_hh],
replace = T 
)

 n.reserve <- min(3, ifelse(
 nrow(idp.frame[COUNTY == idp.breakdown[x, County]]) - idp.breakdown[x, idp_cluster] <
  0,0,
nrow(idp.frame[COUNTY == idp.breakdown[x, County]]) - idp.breakdown[x, idp_cluster]
))
if(n.reserve>0){
reserve.sample <- sample(
  idp.frame[COUNTY == idp.breakdown[x, County] & GEOID9 %nin% main.sample, GEOID9],
  n.reserve,
  prob = idp.frame[COUNTY == idp.breakdown[x, County] & GEOID9 %nin% main.sample, 
  mt12_est_idps_hh],
  replace = F 
)
}else{
reserve.sample <- c()
 }

 data.frame(
 GEOID9 = c(main.sample, reserve.sample),
 sample_type = c(rep("Main sample", idp.breakdown[x, idp_cluster]),
                rep("Reserve",  n.reserve)))
}

idp.sample <- rbindlist(lapply(1:nrow(idp.breakdown), idpSample))

样本数据框

idp.breakdown 前几行

State    County idp_cluster idp_n ret_cluster ret_n hc_cluster hc_n total_cluster total_n
1: Abyei Administrative Area     Abyei          11   143          11   143         11  143            33     429
2:         Central Equatoria      Juba           5    65           3    39          4   52            12     156
3:         Central Equatoria Kajo-Keji           1    13           3    39          8  104            12     156
4:         Central Equatoria    Lainya           4    52           2    26          4   52            10     130
5:         Central Equatoria    Morobo           5    65           3    39          3   39            11     143
6:         Central Equatoria  Terekeka           1    13           1    13         10  130            12     156

idp.frame 前几行

STA_CODE      STATE CTY_CODE COUNTY PAY_CODE BOM_CODE FSNMS_Urban mt12_est_idps_ind mt12_est_idps_hh
1:       71 Upper Nile       01   Renk       01      001       FALSE              1677              304
2:       71 Upper Nile       01   Renk       01      003       FALSE              1104              184
3:       71 Upper Nile       01   Renk       01      004       FALSE               161               31
4:       71 Upper Nile       01   Renk       02      001       FALSE                87               16
5:       71 Upper Nile       01   Renk       02      003       FALSE               796              159
6:       71 Upper Nile       01   Renk       02      006       FALSE               419               82

mt12_est_rets_ind mt12_est_rets_hh    PAYAM      BOMA    GEOID9
1:              5087             1010 Chemmedi Bebineeth 710101001
2:              9446             1892 Chemmedi  Chemmedi 710101003
3:              1766              347 Chemmedi    Labior 710101004
4:              6854             1349    Geger   Gerbena 710102001
5:             26238             5147    Geger  Gospamni 710102003
6:              5967             1195    Geger   Wunthow 710102006
解决方法

这个错误的核心是sample()函数接收的第一个参数无效——要么是空向量,要么不是可抽样的对象。结合你的代码和数据,问题主要出在这几个地方:

  • County名称不匹配:idp.breakdown里的部分County(比如"Abyei")在idp.frame中没有对应数据,导致筛选后返回空向量,无法抽样。
  • %nin%运算符依赖问题:如果未加载data.table或dplyr,这个运算符未定义,会导致筛选未抽样数据时出错。
  • 缺少空值判断:没有提前检查待抽样的数据集是否为空,直接执行sample()触发报错。

可以按以下步骤修复:

  1. 检查并修复County匹配问题
    先找出idp.breakdown中在idp.frame里没有对应数据的County:

    missing_counties <- setdiff(idp.breakdown$County, idp.frame$County)
    print(missing_counties)
    

    针对这些County,要么补充idp.frame的数据,要么在函数中跳过这些行。

  2. 替换%nin%为基础R写法
    用!GEOID9 %in% main.sample替代GEOID9 %nin% main.sample,避免依赖第三方包的运算符:

    # 修正后的预留抽样部分
    reserve.sample <- sample(
      idp.frame[COUNTY == idp.breakdown[x, County] & !GEOID9 %in% main.sample, GEOID9],
      n.reserve,
      prob = idp.frame[COUNTY == idp.breakdown[x, County] & !GEOID9 %in% main.sample, mt12_est_idps_hh],
      replace = F 
    )
    
  3. 在函数中添加空值判断
    在执行抽样前先检查数据集是否为空,避免报错:

    idpSample <- function(x) {
      # 获取当前County的数据集
      county_data <- idp.frame[COUNTY == idp.breakdown[x, County]]
      if(nrow(county_data) == 0) {
        warning(paste("无对应数据的County:", idp.breakdown[x, County]))
        return(data.frame(GEOID9=character(), sample_type=character()))
      }
      
      # 主抽样
      main.sample <- sample(
        county_data$GEOID9,
        idp.breakdown[x, idp_cluster],
        prob = county_data$mt12_est_idps_hh,
        replace = T 
      )
      
      # 计算预留抽样数量
      n.reserve <- min(3, max(0, nrow(county_data) - length(main.sample)))
      reserve.sample <- c()
      if(n.reserve > 0) {
        remaining_data <- county_data[!GEOID9 %in% main.sample]
        if(nrow(remaining_data) >= n.reserve) {
          reserve.sample <- sample(
            remaining_data$GEOID9,
            n.reserve,
            prob = remaining_data$mt12_est_idps_hh,
            replace = F 
          )
        }
      }
      
      # 返回结果
      data.frame(
        GEOID9 = c(main.sample, reserve.sample),
        sample_type = c(rep("Main sample", length(main.sample)),
                       rep("Reserve", length(reserve.sample)))
      )
    }
    
  4. 验证抽样参数
    确认idp.breakdown$idp_cluster都是正整数,避免传入0或负数导致抽样失败。

内容的提问来源于stack exchange,提问作者Emmanuel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 15:24:18