You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于前一行条件替换数据框特定文本的R语言技术问询

国会记录文本分析:参议员指代替换问题

需求说明

  • 对国会记录做文本分析,核心任务是处理参议员之间的指代替换:当参议员用「my friend」「my colleague」「the senator from XX」这类表述指代上一位发言的参议员时,要把这些表述替换成对应参议员的姓名
  • 发言内容按行拆分,每行开头都标注了发言参议员的姓名

试过的方法及遇到的问题

  • if-elseif函数:报错「条件长度大于1,仅使用第一个元素」
  • ifelse函数:没报错,但完全没完成文本替换
  • 带for循环的ifelse函数:直接返回NULL值

测试数据

# 测试环境数据
test_col <- c("Mr. SHELBY (R; Alabama): I acknowledge this is a test.",
              "Mrs. MURRAY (D; Washington): I say to my friend, the senator from Alabama, that they are wrong.",
              "Mr. SHELBY (R; Alabama): I do not agree with my colleague.",
              "Mr. FRIST (R; Tennessee): The senator from Alabama is correct, senator Murray.",
              "Mr. SHELBY (R; Alabama): I thank the majority leader for their support.",
              "Mr. SESSIONS (R; Alabama): I am proud of my junior, the senator from Alabama.",
              "Mr. SHELBY (R; Alabama): To my senior peer, the senator from Alabama, I say great things.")
test_df <- data.frame(test_col)
colnames(test_df) <- c("speeches")

解决方案

要搞定这个替换,得先把每一行的发言者信息提取出来,跟踪上一位发言者的姓名和所属州,再针对性替换指代内容,具体步骤如下:

步骤1:提取发言者关键信息

先从每行发言里拆分出发言者的姓名、所属州:

library(stringr)
library(dplyr)

# 拆分发言者姓名、所属州
test_df <- test_df %>%
  mutate(
    # 提取完整发言者称谓(比如Mr. SHELBY)
    speaker = str_extract(speeches, "^[A-Za-z.\\s]+(?=\\s\\()"),
    # 提取所属州
    state = str_extract(speeches, "(?<=;\\s)[A-Za-z]+(?=\\))"),
    # 提取姓氏(比如SHELBY)
    last_name = str_extract(speaker, "(?<=\\.\\s)[A-Z]+")
  ) %>%
  # 生成上一位发言者的信息列
  mutate(
    prev_last_name = lag(last_name),
    prev_state = lag(state)
  )

步骤2:逐行替换指代内容

用mapply逐行处理,避免向量操作的长度问题,同时定义替换规则:

# 逐行处理替换
test_df$processed_speeches <- mapply(function(speech, prev_last, prev_st) {
  # 先替换通用指代:my friend、my colleague
  updated_speech <- str_replace_all(speech, "my friend|my colleague", prev_last)
  # 替换州指代:比如the senator from Alabama
  if (!is.na(prev_st)) {
    state_pattern <- str_c("the senator from ", prev_st)
    updated_speech <- str_replace_all(updated_speech, state_pattern, prev_last)
  }
  updated_speech
}, test_df$speeches, test_df$prev_last_name, test_df$prev_state)

效果说明

  • 解决了之前ifelse/if-elseif的向量长度不匹配问题,逐行处理更精准
  • 后续可以扩展替换规则,比如添加更多指代短语(比如"my junior"、"my senior peer"),或者覆盖所有州的匹配,只需要修改替换逻辑里的正则表达式即可

内容的提问来源于stack exchange,提问作者Rochelle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 07:57:23