SAS文本处理技术问询:解决compress函数去数字合并文本问题及提取指定关键词前内容
嘿,我来帮你搞定这两个SAS文本处理的问题!咱们一步步拆解解决:
解决SAS文本处理的两个核心问题
问题1:移除数字但保留句子可读性
你说用compress函数会把内容压缩成单个单词,大概率是不小心用了会移除空格的修饰符。其实只要精准指定只移除数字,就能保留原有空格和句子结构:
- 正确用法:
compress(Comments, '0123456789') - 这个函数只会删除0-9的数字字符,空格、字母、特殊符号都会原封不动保留,完全不影响句子可读性。
问题2:提取"documented"之前的所有单词
要精准截取关键词前的内容,咱们可以结合index和substr函数实现:
- 用
index(Comments, 'documented')定位"documented"在文本中的起始位置 - 用
substr(Comments, 1, 起始位置-1)截取该位置之前的文本 - 最后用
trim()去掉末尾多余的空格,得到干净的结果
完整SAS代码示例
下面是针对你提供的示例数据集的完整代码,同时解决两个问题并生成你要的Results列:
/* 创建输入数据集 */ data have; input ID Comments $100.; datalines; 1 increase documented this credit package requires approval 2 new business modification documented ls&f cancelled 3 annual renewal documented this package requires approval ; run; /* 处理文本并生成结果数据集 */ data want; set have; /* 问题1:移除数字(如果存在),保留空格和原结构 */ cleaned_comments = compress(Comments, '0123456789'); /* 问题2:提取"documented"之前的内容 */ if index(cleaned_comments, 'documented') > 0 then do; Results = trim(substr(cleaned_comments, 1, index(cleaned_comments, 'documented') - 1)); end; /* 处理没有"documented"的情况,这里设为空字符串,你可以按需调整 */ else Results = ''; drop cleaned_comments; run; /* 查看输出结果 */ proc print data=want; run;
运行这段代码后,就能得到你期望的输出:
ID Comments Results
1 increase documented this credit package requires approval increase
2 new business modification documented ls&f cancelled new business modification
3 annual renewal documented this package requires approval annual renewal
额外小提示
- 如果文本中"documented"有大小写变化(比如Documented、DOCUMENTED),可以统一大小写后再查找:
index(lowcase(cleaned_comments), 'documented') - 如果文本中有多个"documented",这段代码会提取第一个出现的关键词之前的内容,完全匹配你的示例需求
内容的提问来源于stack exchange,提问作者Learner_new
相关产品推荐
相关产品推荐

