如何用R批量提取Outlook邮件指定属性数据?
问题描述
我们团队每天收到多份带附件的请求邮件,附件是需要成员填写后回复给请求方的表单,邮件响应速度是核心绩效指标之一。我用Outlook文件夹专门存放这些请求与回复邮件,目前使用Microsoft365R包,但无法批量提取所需信息。
我尝试过的代码:
library(Microsoft365R) # 认证Microsoft 365(需完成交互式登录) get_business_onedrive() # 会触发登录提示 outlook <- get_business_outlook() # 指定文件夹名称(替换为实际文件夹名) folder_name <- "PRA Checklists" folder <- outlook$get_folder(folder_name) # 获取文件夹内的邮件 emails <- folder$list_emails(n = 100000) # 查看emails对象结构 str(emails) ls(emails[[1]]) lapply(emails, ls)
emails是一个环境列表,每个环境中的properties包含我需要的数据,但我不知道如何批量整理为数据框。需要提取的属性包括:c("receivedDateTime", "sentDateTime", "hasAttachments", "subject", "bodyPreview", "conversationId")
我已成功连接Outlook并获取环境对象列表,但高效提取数据失败,尝试的for循环也无效:
for (i in 1:length(emails)) { for (j in 1:length(emails[[i]]$properties)){ list[[i]] <- emails[[i]]$properties[[j]] } }
请求提供高效提取这些数据到数据框的方法。
解决方案
可以通过两种方式实现高效提取,分别是依赖purrr的简洁写法,以及基础R原生写法:
方法一:使用purrr包(推荐)
借助映射函数批量处理列表,代码更简洁且不易出错:
library(Microsoft365R) library(purrr) library(dplyr) # 定义需要提取的目标属性 target_fields <- c("receivedDateTime", "sentDateTime", "hasAttachments", "subject", "bodyPreview", "conversationId") # 遍历邮件列表,提取属性并合并为数据框 email_data <- map_dfr(emails, function(email) { # 提取每个目标字段,不存在的字段自动返回NA map(target_fields, ~ email$properties[[.x]]) %>% set_names(target_fields) %>% as.data.frame() }) # 验证结果 str(email_data) head(email_data)
代码说明
map_dfr:遍历emails列表,将每个邮件的提取结果行绑定为一个完整数据框- 内层
map:针对每个目标字段从properties中取值,字段不存在时返回NA,避免报错 set_names:为提取的向量设置字段名,保证数据框列名与目标属性一致
方法二:基础R原生写法
如果不想额外加载包,可使用基础R实现:
library(Microsoft365R) # 定义需要提取的目标属性 target_fields <- c("receivedDateTime", "sentDateTime", "hasAttachments", "subject", "bodyPreview", "conversationId") # 初始化空列表存储单条邮件数据 email_list <- vector("list", length(emails)) # 循环提取每个邮件的目标属性 for (i in seq_along(emails)) { props <- emails[[i]]$properties # 提取字段,不存在则返回NA email_list[[i]] <- data.frame( receivedDateTime = props[["receivedDateTime"]], sentDateTime = props[["sentDateTime"]], hasAttachments = props[["hasAttachments"]], subject = props[["subject"]], bodyPreview = props[["bodyPreview"]], conversationId = props[["conversationId"]], stringsAsFactors = FALSE ) } # 合并为完整数据框 email_data <- do.call(rbind, email_list) # 验证结果 str(email_data) head(email_data)
代码说明
- 提前初始化对应长度的空列表,避免循环中动态扩容,提升运行效率
- 直接指定字段提取,不存在的字段会返回
NULL,转换为数据框时自动转为NA do.call(rbind, email_list)将列表中的所有数据框合并为一个大的数据框
内容的提问来源于stack exchange,提问作者Matt Simonson
相关产品推荐
相关产品推荐

