基于前一行条件替换数据框特定文本的R语言技术问询
国会记录文本分析:参议员指代替换问题
需求说明
- 对国会记录做文本分析,核心任务是处理参议员之间的指代替换:当参议员用「my friend」「my colleague」「the senator from XX」这类表述指代上一位发言的参议员时,要把这些表述替换成对应参议员的姓名
- 发言内容按行拆分,每行开头都标注了发言参议员的姓名
试过的方法及遇到的问题
- if-elseif函数:报错「条件长度大于1,仅使用第一个元素」
- ifelse函数:没报错,但完全没完成文本替换
- 带for循环的ifelse函数:直接返回NULL值
测试数据
# 测试环境数据 test_col <- c("Mr. SHELBY (R; Alabama): I acknowledge this is a test.", "Mrs. MURRAY (D; Washington): I say to my friend, the senator from Alabama, that they are wrong.", "Mr. SHELBY (R; Alabama): I do not agree with my colleague.", "Mr. FRIST (R; Tennessee): The senator from Alabama is correct, senator Murray.", "Mr. SHELBY (R; Alabama): I thank the majority leader for their support.", "Mr. SESSIONS (R; Alabama): I am proud of my junior, the senator from Alabama.", "Mr. SHELBY (R; Alabama): To my senior peer, the senator from Alabama, I say great things.") test_df <- data.frame(test_col) colnames(test_df) <- c("speeches")
解决方案
要搞定这个替换,得先把每一行的发言者信息提取出来,跟踪上一位发言者的姓名和所属州,再针对性替换指代内容,具体步骤如下:
步骤1:提取发言者关键信息
先从每行发言里拆分出发言者的姓名、所属州:
library(stringr) library(dplyr) # 拆分发言者姓名、所属州 test_df <- test_df %>% mutate( # 提取完整发言者称谓(比如Mr. SHELBY) speaker = str_extract(speeches, "^[A-Za-z.\\s]+(?=\\s\\()"), # 提取所属州 state = str_extract(speeches, "(?<=;\\s)[A-Za-z]+(?=\\))"), # 提取姓氏(比如SHELBY) last_name = str_extract(speaker, "(?<=\\.\\s)[A-Z]+") ) %>% # 生成上一位发言者的信息列 mutate( prev_last_name = lag(last_name), prev_state = lag(state) )
步骤2:逐行替换指代内容
用mapply逐行处理,避免向量操作的长度问题,同时定义替换规则:
# 逐行处理替换 test_df$processed_speeches <- mapply(function(speech, prev_last, prev_st) { # 先替换通用指代:my friend、my colleague updated_speech <- str_replace_all(speech, "my friend|my colleague", prev_last) # 替换州指代:比如the senator from Alabama if (!is.na(prev_st)) { state_pattern <- str_c("the senator from ", prev_st) updated_speech <- str_replace_all(updated_speech, state_pattern, prev_last) } updated_speech }, test_df$speeches, test_df$prev_last_name, test_df$prev_state)
效果说明
- 解决了之前ifelse/if-elseif的向量长度不匹配问题,逐行处理更精准
- 后续可以扩展替换规则,比如添加更多指代短语(比如"my junior"、"my senior peer"),或者覆盖所有州的匹配,只需要修改替换逻辑里的正则表达式即可
内容的提问来源于stack exchange,提问作者Rochelle
相关产品推荐
相关产品推荐

