You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中利用通用职位语料库统计完整职位头衔的出现频率?

问题描述

我有一组包含简短职位描述和职位头衔的数据集,想要分析其中职位头衔的出现频率。

数据示例

姓名职位描述
JohnDigital Product Lead
JaneAccount Manager, Head of Workplace Experience
BillExecutive Assistant at EY
TomSenior Recruiter, People Experience, Talent Branding #EX #DX

职位字段里存在一些无关字符串,我之前用简单NLP方法生成了词汇及其语料库频率的数据框,所用R代码如下:

text <- df$job

docs <- VCorpus(VectorSource(text))

docs <- docs %>%
  tm_map(removeNumbers) %>%
  tm_map(removePunctuation) %>%
  tm_map(stripWhitespace)

docs <- tm_map(docs, content_transformer(tolower))
docs <- tm_map(docs, removeWords, stopwords("english"))

dtm <- TermDocumentMatrix(docs) 
matrix <- as.matrix(dtm) 
words <- sort(rowSums(matrix),decreasing=TRUE) 
JobFreq <- data.frame(word = names(words),freq=words)

返回的数据示例如下:

词汇频率
Senior1656
Analyst798

返回的数据接近需求,但多数职位头衔是多词组合,现在的处理把它们拆成了单个词汇,丢失了上下文。请问能不能导入通用职位头衔语料库,让数据中匹配语料库的多词字符串不被分词,作为整体处理?如果可以,怎么用R实现?


解决方案

完全可以通过自定义多词职位头衔词典,让NLP工具把这些多词组合当作单个"词汇"处理,避免被拆分。下面提供两种在R中实现的实用思路:

方法一:基于tm包的自定义短语匹配

  1. 准备职位头衔语料库
    先整理好你的多词职位头衔列表,统一转成小写(和后续文本预处理格式对齐):

    job_titles <- c("digital product lead", "account manager", "executive assistant", 
                    "senior recruiter", "head of workplace experience")
    
  2. 编写自定义分词函数
    优先匹配多词头衔,再处理剩余文本:

    # 把职位头衔转成正则表达式,确保匹配完整短语
    title_pattern <- paste0("\\b", job_titles, "\\b", collapse = "|")
    
    custom_tokenizer <- function(x) {
      # 提取所有匹配的多词头衔
      matches <- str_extract_all(x, title_pattern)[[1]]
      # 移除已匹配的内容,处理剩余文本
      remaining_text <- str_remove_all(x, title_pattern)
      # 对剩余文本做常规分词
      remaining_tokens <- unlist(str_split(remaining_text, "\\s+"))
      # 过滤空字符串,合并结果
      c(matches, remaining_tokens)[c(matches, remaining_tokens) != ""]
    }
    
  3. 替换tm默认分词器并计算频率

    # 文本预处理(只做必要清洗,跳过默认分词步骤)
    docs <- VCorpus(VectorSource(text))
    docs <- docs %>%
      tm_map(content_transformer(tolower)) %>%
      tm_map(removeNumbers) %>%
      tm_map(removePunctuation) %>%
      tm_map(stripWhitespace)
    
    # 用自定义分词器构建词项-文档矩阵
    dtm <- TermDocumentMatrix(docs, control = list(tokenize = custom_tokenizer))
    
    # 计算频率并整理成数据框
    matrix <- as.matrix(dtm)
    words <- sort(rowSums(matrix), decreasing = TRUE)
    JobFreq <- data.frame(word = names(words), freq = words)
    

方法二:使用quanteda包(更简洁的短语匹配)

quanteda对多词短语的支持更原生,步骤更简洁:

  1. 安装并加载包

    install.packages("quanteda")
    library(quanteda)
    
  2. 创建自定义职位词典

    job_dict <- dictionary(list(job_titles = job_titles))
    
  3. 处理文本并统计频率

    # 创建语料库并预处理
    corp <- corpus(text) %>%
      tokens(remove_numbers = TRUE, remove_punct = TRUE, remove_symbols = TRUE) %>%
      tokens_tolower() %>%
      tokens_remove(stopwords("english"))
    
    # 将匹配的多词短语合并为单个token
    corp_phrases <- tokens_compound(corp, pattern = phrase(job_titles))
    
    # 统计频率并排序
    job_freq <- dfm(corp_phrases) %>%
      textstat_frequency() %>%
      select(feature, frequency) %>%
      arrange(desc(frequency))
    

注意事项

  • 职位头衔语料库越全面,匹配效果越好,可以从公开职业分类数据集(比如ONET职业库)中提取整理。
  • 预处理时要保证文本和词典格式一致(比如统一小写、去除特殊字符),避免匹配失败。
  • 如果职位头衔有变体(比如"senior recruiter"和"sr. recruiter"),可以在正则表达式中加入变体规则,或者先对文本做标准化处理。

内容的提问来源于stack exchange,提问作者xanban

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 14:40:18