R语言解析市议会动议长字符串并提取结构化字段的方法咨询
解析动议文本的实现方案
R实现方法
你可以直接用tidyverse生态的stringr做正则匹配、tidyr做列结构转换,不需要额外安装小众包,完整实现代码如下:
# 未安装依赖先执行:install.packages("tidyverse") library(tidyverse) # 匹配规则:动议内容(数字点+内容) + 提交人(Councillor/Mayor + 人名) pattern <- "(\\d+\\. .*?)\\s+((?:Councillor|Mayor)\\s+[A-Za-z]+)" dat_processed <- dat %>% # 提取每行所有匹配的动议、提交人对 mutate(matches = str_match_all(text, pattern)) %>% rowwise() %>% mutate( motions = list(matches[,2]), submitters = list(matches[,3]) ) %>% ungroup() %>% # 列表展开为多列 unnest_wider(c(motions, submitters), names_sep = "_") %>% # 按要求重命名列 rename_with(~str_replace(., "motions_(\\d+)", "Motion\\1"), starts_with("motions_")) %>% rename_with(~str_replace(., "submitters_(\\d+)", "Motion\\1_Submitter"), starts_with("submitters_")) %>% # 空值替换为NA字符串 mutate(across(starts_with("Motion"), ~ifelse(is.na(.), "NA", .)))
如果实际场景中人名存在连字符、空格等特殊字符,只需要调整正则中匹配人名的部分即可适配。
Python实现难度对比
Python实现逻辑和R完全一致,代码量、复杂度和R没有明显差异,不存在更简便的说法,你可以根据自己常用的技术栈选择即可,Python参考代码如下:
import pandas as pd import re pattern = r"(\d+\. .*?)\s+((?:Councillor|Mayor)\s+[A-Za-z]+)" def parse_text(text): matches = re.findall(pattern, text) res = {} for idx, (motion, submitter) in enumerate(matches, 1): res[f"Motion{idx}"] = motion res[f"Motion{idx}_Submitter"] = submitter return res parsed_cols = dat['text'].apply(parse_text).apply(pd.Series) dat_processed = pd.concat([dat, parsed_cols], axis=1).fillna("NA")
内容的提问来源于stack exchange,提问作者scotiaboy
相关产品推荐
相关产品推荐

