学校系统数据处理需求:将HAVE列按大写字母规则拆分为WANT1和WANT2
数据拆分解决方案
需求说明
将HAVE列数据拆分为WANT1和WANT2两列,规则如下:
- WANT1:提取开头连续大写字母中的前n-1个(n为开头连续大写字母的总数)
- WANT2:提取剩余的所有字符
示例数据
HAVE WANT1 WANT2 CLStephen Five CL Stephen Five RTQQuent Lou X RTQ Quent Lou X
实现方法
Python 代码示例
利用正则表达式匹配开头连续大写字母,再按规则拆分:
import re def split_have(input_str): cap_match = re.match(r'([A-Z]+)(.*)', input_str) if cap_match: all_caps, remaining = cap_match.groups() want1 = all_caps[:-1] # 拼接最后一个大写字母和剩余内容,去除多余空格 want2 = f"{all_caps[-1]}{remaining}".strip() return want1, want2 return "", input_str # 测试示例数据 test_cases = ["CLStephen Five", "RTQQuent Lou X"] for case in test_cases: w1, w2 = split_have(case) print(f"{case:<20} {w1:<6} {w2}")
执行后输出:
CLStephen Five CL Stephen Five RTQQuent Lou X RTQ Quent Lou X
SAS 代码示例
使用PRX系列函数实现拆分逻辑:
data want_table; set have_table; regex = prxparse('/^([A-Z]+)(.*)/'); if prxmatch(regex, HAVE) then do; cap_str = prxposn(regex, 1, HAVE); rest_str = prxposn(regex, 2, HAVE); WANT1 = substr(cap_str, 1, length(cap_str)-1); WANT2 = strip(cats(substr(cap_str, length(cap_str)), rest_str)); end; drop regex cap_str rest_str; run;
内容的提问来源于stack exchange,提问作者bvowe
相关产品推荐
相关产品推荐

