You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言解析市议会动议长字符串并提取结构化字段的方法咨询

解析动议文本的实现方案

R实现方法

你可以直接用tidyverse生态的stringr做正则匹配、tidyr做列结构转换,不需要额外安装小众包,完整实现代码如下:

# 未安装依赖先执行:install.packages("tidyverse")
library(tidyverse)

# 匹配规则:动议内容(数字点+内容) + 提交人(Councillor/Mayor + 人名)
pattern <- "(\\d+\\. .*?)\\s+((?:Councillor|Mayor)\\s+[A-Za-z]+)"

dat_processed <- dat %>%
  # 提取每行所有匹配的动议、提交人对
  mutate(matches = str_match_all(text, pattern)) %>%
  rowwise() %>%
  mutate(
    motions = list(matches[,2]),
    submitters = list(matches[,3])
  ) %>%
  ungroup() %>%
  # 列表展开为多列
  unnest_wider(c(motions, submitters), names_sep = "_") %>%
  # 按要求重命名列
  rename_with(~str_replace(., "motions_(\\d+)", "Motion\\1"), starts_with("motions_")) %>%
  rename_with(~str_replace(., "submitters_(\\d+)", "Motion\\1_Submitter"), starts_with("submitters_")) %>%
  # 空值替换为NA字符串
  mutate(across(starts_with("Motion"), ~ifelse(is.na(.), "NA", .)))

如果实际场景中人名存在连字符、空格等特殊字符,只需要调整正则中匹配人名的部分即可适配。

Python实现难度对比

Python实现逻辑和R完全一致,代码量、复杂度和R没有明显差异,不存在更简便的说法,你可以根据自己常用的技术栈选择即可,Python参考代码如下:

import pandas as pd
import re

pattern = r"(\d+\. .*?)\s+((?:Councillor|Mayor)\s+[A-Za-z]+)"

def parse_text(text):
    matches = re.findall(pattern, text)
    res = {}
    for idx, (motion, submitter) in enumerate(matches, 1):
        res[f"Motion{idx}"] = motion
        res[f"Motion{idx}_Submitter"] = submitter
    return res

parsed_cols = dat['text'].apply(parse_text).apply(pd.Series)
dat_processed = pd.concat([dat, parsed_cols], axis=1).fillna("NA")

内容的提问来源于stack exchange,提问作者scotiaboy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 06:27:01