拆分单列数据统计出版物数量与总引用数(R/Stata/Python)
解决方案
R 实现
借助tidyverse工具集的stringr包即可完成处理:
library(tidyverse) # 构造示例数据 df <- tibble( raw_data = c("21070808(136)|19995886(87)|21280165(66)", "20226255(57)|21440646(54)") ) # 生成目标列 df_processed <- df %>% mutate( # 统计出版物数量:|的个数加1 pub_count = str_count(raw_data, "\\|") + 1, # 提取所有括号内的引用数并求和 total_citations = map_int(raw_data, ~{ str_extract_all(.x, "\\((\\d+)\\)")[[1]] %>% str_remove_all("[()]") %>% as.integer() %>% sum() }) ) %>% select(pub_count, total_citations) # 查看结果 print(df_processed)
输出结果:
# A tibble: 2 × 2 pub_count total_citations <int> <int> 1 3 289 2 2 111
Stata 实现
通过正则表达式和循环处理,两种方式可选:
方式1:直接提取所有引用数
* 示例数据 clear input strL raw_data "21070808(136)|19995886(87)|21280165(66)" "20226255(57)|21440646(54)" end * 计算出版物数量 gen pub_count = regexrcount(raw_data, "\|") + 1 * 提取并累加引用数 gen total_citations = 0 tempvar cit_str forvalues i = 1/100 { // 假设每行最多100个出版物,可按需调整 capture regexs(raw_data, "\((\d+)\)", `i', `cit_str') if _rc break replace total_citations = total_citations + real(`cit_str') } * 查看结果 list pub_count total_citations
方式2:拆分后逐个处理
split raw_data, parse("|") gen pub_count = r(nvars) gen total_citations = 0 forvalues i = 1/`r(nvars)' { replace total_citations = total_citations + real(ustrregexs(1)) if ustrregexm(raw_data`i', "\((\d+)\)") } drop raw_data1-raw_data`r(nvars)'
Python 实现
用pandas结合正则表达式快速处理:
import pandas as pd import re # 示例数据 raw_data = [ "21070808(136)|19995886(87)|21280165(66)", "20226255(57)|21440646(54)" ] df = pd.DataFrame({"raw_data": raw_data}) # 定义行处理函数 def calc_stats(row): entries = row.split("|") pub_count = len(entries) # 提取所有括号内的数字并求和 citations = [int(re.search(r"\((\d+)\)", entry).group(1)) for entry in entries] return pd.Series([pub_count, sum(citations)]) # 生成目标列 df[["pub_count", "total_citations"]] = df["raw_data"].apply(calc_stats) # 输出结果 print(df[["pub_count", "total_citations"]])
输出结果:
pub_count total_citations 0 3 289 1 2 111
内容的提问来源于stack exchange,提问作者ClaraJ
相关产品推荐
相关产品推荐

