You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R批量提取Outlook邮件指定属性数据?

问题描述

我们团队每天收到多份带附件的请求邮件,附件是需要成员填写后回复给请求方的表单,邮件响应速度是核心绩效指标之一。我用Outlook文件夹专门存放这些请求与回复邮件,目前使用Microsoft365R包,但无法批量提取所需信息。

我尝试过的代码:

library(Microsoft365R)
# 认证Microsoft 365(需完成交互式登录)
get_business_onedrive()  # 会触发登录提示
outlook <- get_business_outlook()

# 指定文件夹名称(替换为实际文件夹名)
folder_name <- "PRA Checklists"
folder <- outlook$get_folder(folder_name)

# 获取文件夹内的邮件
emails <- folder$list_emails(n = 100000)

# 查看emails对象结构
str(emails)

ls(emails[[1]])
lapply(emails, ls)

emails是一个环境列表,每个环境中的properties包含我需要的数据,但我不知道如何批量整理为数据框。需要提取的属性包括:c("receivedDateTime", "sentDateTime", "hasAttachments", "subject", "bodyPreview", "conversationId")

我已成功连接Outlook并获取环境对象列表,但高效提取数据失败,尝试的for循环也无效:

for (i in 1:length(emails)) { 
  for (j in 1:length(emails[[i]]$properties)){
    list[[i]] <- emails[[i]]$properties[[j]]
  }  
  }

请求提供高效提取这些数据到数据框的方法。


解决方案

可以通过两种方式实现高效提取,分别是依赖purrr的简洁写法,以及基础R原生写法:

方法一:使用purrr包(推荐)

借助映射函数批量处理列表,代码更简洁且不易出错:

library(Microsoft365R)
library(purrr)
library(dplyr)

# 定义需要提取的目标属性
target_fields <- c("receivedDateTime", "sentDateTime", "hasAttachments", "subject", "bodyPreview", "conversationId")

# 遍历邮件列表,提取属性并合并为数据框
email_data <- map_dfr(emails, function(email) {
  # 提取每个目标字段,不存在的字段自动返回NA
  map(target_fields, ~ email$properties[[.x]]) %>%
    set_names(target_fields) %>%
    as.data.frame()
})

# 验证结果
str(email_data)
head(email_data)

代码说明

  • map_dfr:遍历emails列表,将每个邮件的提取结果行绑定为一个完整数据框
  • 内层map:针对每个目标字段从properties中取值,字段不存在时返回NA,避免报错
  • set_names:为提取的向量设置字段名,保证数据框列名与目标属性一致

方法二:基础R原生写法

如果不想额外加载包,可使用基础R实现:

library(Microsoft365R)

# 定义需要提取的目标属性
target_fields <- c("receivedDateTime", "sentDateTime", "hasAttachments", "subject", "bodyPreview", "conversationId")

# 初始化空列表存储单条邮件数据
email_list <- vector("list", length(emails))

# 循环提取每个邮件的目标属性
for (i in seq_along(emails)) {
  props <- emails[[i]]$properties
  # 提取字段,不存在则返回NA
  email_list[[i]] <- data.frame(
    receivedDateTime = props[["receivedDateTime"]],
    sentDateTime = props[["sentDateTime"]],
    hasAttachments = props[["hasAttachments"]],
    subject = props[["subject"]],
    bodyPreview = props[["bodyPreview"]],
    conversationId = props[["conversationId"]],
    stringsAsFactors = FALSE
  )
}

# 合并为完整数据框
email_data <- do.call(rbind, email_list)

# 验证结果
str(email_data)
head(email_data)

代码说明

  • 提前初始化对应长度的空列表,避免循环中动态扩容,提升运行效率
  • 直接指定字段提取,不存在的字段会返回NULL,转换为数据框时自动转为NA
  • do.call(rbind, email_list)将列表中的所有数据框合并为一个大的数据框

内容的提问来源于stack exchange,提问作者Matt Simonson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 15:22:03