从standard-input读取行的性能问题及优化方案咨询
SBCL处理标准输入的性能瓶颈优化方案
问题背景
需要处理100GB格式特殊的日志,已完成解析和CLI初始开发。测试1GB数据时耗时约1分钟;做了一项sanity check:仅将standard-input内容复制到standard-out,结果显示大部分时间消耗在读取环节,而Python完成相同操作仅需几秒。此前Common Lisp在复杂数值计算上性能优于Python,现寻求解决或规避该性能问题的方案。
测试环境与数据
生成测试数据命令
yes "$(printf 'A%.0s' {1..10})" | head -c 1G > sample.txt
测试代码
Common Lisp代码(标准输入复制)
#!/usr/bin/env -S sbcl --script (defun main () (loop for line = (read-line *standard-input* nil) while line do (write-string line) (write-char #\NEWLINE))) (eval-when (:execute) (main))
Python代码(标准输入复制)
#!/usr/bin/env python3 import sys def main(): for line in sys.stdin: sys.stdout.write(line) if __name__ == "__main__": main()
运行结果
time cat data/sample.txt | ./test.lisp > result_lisp.txt # real 1m59.719s # user 0m47.231s # sys 1m13.008s time cat data/sample.txt | ./test.py > result_python.txt # real 0m9.557s # user 0m7.688s # sys 0m2.144s
版本信息
- SBCL版本:SBCL 2.2.9.debian
- Python版本:Python 3.13.3
性能分析
仅读取标准输入的性能分析代码
#!/usr/bin/env -S sbcl --script (require :sb-sprof) (defun main () (loop for line = (read-line *standard-input* nil) while line)) (eval-when (:execute) (sb-sprof:with-profiling (:max-samples 100000 :sample-interval 0.00001 :report :graph) (main)))
标准输入读取分析结果
大部分时间消耗在UTF-8字符转换上:
Self Total Cumul Nr Count % Count % Count % Calls Function ------------------------------------------------------------------------ 1 12930 49.2 13191 50.2 12930 49.2 - SB-IMPL::INPUT-CHAR/UTF-8 2 9724 37.0 22523 85.7 22654 86.2 - (LAMBDA (&REST REST) :IN SB-IMPL::GET-EXTERNAL-FORMAT) 3 2517 9.6 25834 98.3 25171 95.8 - READ-LINE 4 146 0.6 146 0.6 25317 96.3 - foreign function pthread_sigmask 5 26 0.1 26257 99.9 25343 96.4 - MAIN 6 26 0.1 26 0.1 25369 96.5 - RESTORE-YMM 7 6 0.0 262 1.0 25375 96.5 - SB-IMPL::REFILL-INPUT-BUFFER 8 5 0.0 214 0.8 25380 96.6 - (FLET "WITHOUT-INTERRUPTS-BODY-2" :IN SB-IMPL::REFILL-INPUT-BUFFER) 9 5 0.0 5 0.0 25385 96.6 - SAVE-YMM 10 4 0.0 617 2.3 25389 96.6 - ALLOC-TRAMP
直接读取文件的性能分析代码
(require :sb-sprof) (defun main (str) (loop for line = (read-line str nil) while line)) (eval-when (:execute) (sb-sprof:with-profiling (:max-samples 100000 :sample-interval 0.00001 :report :graph) (with-open-file (str "sample.txt") (main str))))
直接读文件的分析结果
速度大幅提升,调用了更高效的UTF-8处理函数:
Self Total Cumul Nr Count % Count % Count % Calls Function ------------------------------------------------------------------------ 1 2133 38.7 2395 43.4 2133 38.7 - SB-IMPL::FD-STREAM-READ-N-CHARACTERS/UTF-8 2 763 13.8 5212 94.5 2896 52.5 - SB-IMPL::ANSI-STREAM-READ-LINE-FROM-FRC-BUFFER 3 725 13.1 725 13.1 3621 65.6 - SB-KERNEL:UB32-BASH-COPY 4 435 7.9 1882 34.1 4056 73.5 - (LABELS SB-IMPL::BUILD-RESULT :IN SB-IMPL::ANSI-STREAM-READ-LINE-FROM-FRC-BUFFER) 5 261 4.7 261 4.7 4317 78.2 - READ-LINE 6 119 2.2 119 2.2 4436 80.4 - foreign function pthread_sigmask 7 33 0.6 5516 100.0 4469 81.0 - MAIN 8 16 0.3 2413 43.7 4485 81.3 - SB-INT:FAST-READ-CHAR-REFILL 9 15 0.3 15 0.3 4500 81.6 - RESTORE-YMM 10 11 0.2 11 0.2 4511 81.8 - SAVE-YMM
优化方案
1. 重新配置标准输入的外部格式与缓冲区
如果日志是纯ASCII格式,直接指定外部格式为:ascii并开启全缓冲,避免逐字符UTF-8转换的开销:
#!/usr/bin/env -S sbcl --script (defun main () (let ((input-stream (sb-ext:make-fd-stream 0 :external-format :ascii :buffering :full :buffer-size (* 1024 1024)))) ; 1MB缓冲区 (loop for line = (read-line input-stream nil) while line do (write-string line) (write-char #\NEWLINE)))) (eval-when (:execute) (main))
即使日志是UTF-8,指定:utf-8并加大缓冲区也能减少系统调用和转换次数。
2. 使用二进制流批量处理
绕过字符流的逐行转换,直接读取二进制块,再批量转换为字符串并分割行,适合大文件处理:
#!/usr/bin/env -S sbcl --script (defun process-binary-stream (in out) (let* ((buffer-size (* 16 1024)) ; 16KB缓冲区 (byte-buffer (make-array buffer-size :element-type '(unsigned-byte 8))) (char-buffer (make-array buffer-size :element-type 'character)) (remaining "")) (loop for bytes-read = (read-sequence byte-buffer in) while (> bytes-read 0) do (let* ((current-string (concatenate 'string remaining (sb-ext:octets-to-string byte-buffer :start 0 :end bytes-read :external-format :ascii))) (lines (split-sequence:split-sequence #\Newline current-string))) (setf remaining (car (last lines))) (dolist (line (butlast lines)) (write-string line out) (write-char #\Newline out)))) (unless (string= remaining "") (write-string remaining out)))) (defun main () (process-binary-stream *standard-input* *standard-output*)) (eval-when (:execute) (require :split-sequence) ; 需要先安装split-sequence库 (main))
注意:需先通过ql:quickload :split-sequence安装split-sequence库,或手动实现行分割逻辑。
3. 升级SBCL版本
你使用的SBCL 2.2.9.debian是较旧的发行版打包版本,最新的SBCL(如2.4.x系列)对标准输入流和字符转换有性能优化,升级后可能直接解决该瓶颈。
4. 支持文件参数与管道模式兼顾
修改脚本,优先处理命令行传入的文件名(直接读文件性能更高),无参数时再读取标准输入,兼顾Shell脚本友好性和性能:
#!/usr/bin/env -S sbcl --script (defun process-stream (stream) (loop for line = (read-line stream nil) while line do (write-string line) (write-char #\Newline))) (defun main () (let ((args (cdr sb-ext:*posix-argv*))) (if args (dolist (filename args) (with-open-file (stream filename :external-format :ascii :buffering :full) (process-stream stream))) (let ((input-stream (sb-ext:make-fd-stream 0 :external-format :ascii :buffering :full))) (process-stream input-stream))))) (eval-when (:execute) (main))
内容的提问来源于stack exchange,提问作者nemo nemo
相关产品推荐
相关产品推荐

