You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在CLIPS中实现文本分词及特定规则处理的函数方法?

在CLIPS中实现文本分词及自定义处理函数

CLIPS本身没有内置的文本分词功能,不过我们可以用它的字符串处理工具手动实现,同时把你需要的所有功能整合到一个自定义函数里。下面是完整的实现步骤和代码:

1. 先实现核心辅助函数

首先我们需要两个辅助函数:一个用来分割字符串(分词),另一个用来按规则转换单词。

分词辅助函数

这个函数会把输入的文本按空格分割成单词列表,同时简单清理掉单词首尾的标点(比如逗号、句号):

(defun split-text (input)
  (if (eq (str-length input) 0)
      then ()
      else
      (bind (space-pos) (str-index " " input))
      (if (eq space-pos nil)
          then
          ;; 清理单个单词的标点
          (list (strip-punctuation input))
          else
          (bind (word) (str-substring 1 (- space-pos 1) input))
          (cons (strip-punctuation word)
                (split-text (str-substring (+ space-pos 1) (str-length input) input))))))

(defun strip-punctuation (word)
  ;; 移除单词首尾的常见标点
  (bind (cleaned) word)
  (while (member (str-substring 1 1 cleaned) '(#\, #\. #\! #\?))
    (bind (cleaned) (str-substring 2 (str-length cleaned) cleaned)))
  (while (member (str-substring (str-length cleaned) (str-length cleaned) cleaned) '(#\, #\. #\! #\?))
    (bind (cleaned) (str-substring 1 (- (str-length cleaned) 1) cleaned)))
  cleaned)

单词转换辅助函数

这个函数会检查单词长度,奇数长度时移除中间字符:

(defun transform-word (word)
  (bind (len) (str-length word))
  (if (oddp len)
      then
      (bind (mid-pos) (+ (floor len 2) 1))
      ;; 拼接中间字符前后的部分
      (str-cat (str-substring 1 (- mid-pos 1) word)
               (str-substring (+ mid-pos 1) len word))
      else
      word))

2. 主处理函数

接下来是整合所有功能的主函数:它会完成分词、转换单词、筛选不同单词,同时输出到屏幕和文件:

(defun process-and-output (input-text output-file-path)
  ;; 1. 分词得到单词列表
  (bind (word-list) (split-text input-text))
  (if (eq (length word-list) 0)
      then
      (printout t "输入文本为空!" crlf)
      (return))

  ;; 2. 转换所有单词
  (bind (transformed-list) (mapcar transform-word word-list))
  (bind (first-word) (nth$ 1 transformed-list))

  ;; 3. 筛选出和第一个单词不同的转换后单词
  (bind (filtered-list)
        (filter (lambda (w) (not (eq w first-word))) transformed-list))

  ;; 4. 输出到屏幕
  (printout t "处理结果:" crlf)
  (printout t "原始单词列表:" word-list crlf)
  (printout t "转换后单词列表:" transformed-list crlf)
  (printout t "与第一个单词不同的转换后单词:" filtered-list crlf)

  ;; 5. 写入文件
  (bind (out-file) (open output-file-path "w"))
  (if (eq out-file nil)
      then
      (printout t "无法打开输出文件:" output-file-path crlf)
      (return))
  (printout out-file "处理结果:" crlf)
  (printout out-file "原始单词列表:" word-list crlf)
  (printout out-file "转换后单词列表:" transformed-list crlf)
  (printout out-file "与第一个单词不同的转换后单词:" filtered-list crlf)
  (close out-file)
  (printout t "结果已写入文件:" output-file-path crlf))

3. 测试示例

你可以这样调用这个函数来测试:

(process-and-output "Hello world this is a test sentence example" "output.txt")

运行后,屏幕会输出处理后的结果,同时在当前目录生成output.txt文件,里面包含相同的内容。比如:

  • 原始单词"Hello"长度5(奇数),转换后变成"Hllo"(移除中间的'l')
  • 单词"world"长度5,转换后变成"wold"
  • 最终会筛选出所有不是"Hllo"的转换后单词

注意事项

  • 这个分词逻辑是按空格分割,如果你的文本有其他分隔符(比如制表符),可以修改split-text函数里的分隔符判断
  • strip-punctuation函数只处理了常见的标点,如果需要处理更多符号,可以扩展里面的字符列表
  • CLIPS的字符串索引是从1开始的,所以在使用str-substring时要注意起始和结束位置

内容的提问来源于stack exchange,提问作者Вадим Мороз

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:34:07