在Tcl中使用binary scan读取32位无符号整数遇阻
浏览器Native Messaging与Tcl交互的字节读取问题
我确定自己犯了低级错误,但始终无法解决这个问题:我正在使用浏览器扩展的Native-messaging API将JSON字符串传递给Tcl脚本,该功能在C语言中使用uint32_t可正常实现,但在Tcl中无法正确读取消息前缀的32位无符号整数。我尝试过在binary scan命令中使用nu、Iu和iu格式,写入$len到文件确认有读取内容,但程序似乎卡在coread过程等待读取超出消息实际长度的字节。
消息包含希伯来文、希腊文等多字节字符是否会影响?
后续补充:我确定是多字节字符导致的问题,因为无此类字符时代码可正常运行。但为什么会这样?浏览器和Tcl不都是按字节统计消息长度吗?目前Tcl解析出的长度总是大于实际读取的字节数。
文档说明
每条消息通过JSON序列化、UTF-8编码,前缀为一个按原生字节序存储的无符号32位整数,表示消息长度。
C语言监听stdin代码
int listen( void ) { uint32_t msg_len; while ( fread( &msg_len, sizeof msg_len, 1, stdin ) == 1 ) { char *buf = malloc( msg_len ); if ( !buf ) { } else if ( fread( buf, sizeof *buf, msg_len, stdin ) != msg_len ) { } else { // Read the full messsage. } fflush( stdout ); free( buf ); } return 0; }
尝试的Tcl代码
proc coread {reqBytes} { set ::extMsg {} set remBytes $reqBytes while {![chan eof stdin]} { yield append ::extMsg [set data [read stdin $remBytes]] set remBytes [expr {$remBytes - [string length $data]}] if {$remBytes == 0} { return $reqBytes } } throw {COREAD EOF} "Unexpected EOF" } proc ExtRead {} { chan event stdin readable coro_stdin while {1} { if { [coread 4] != 4 || [binary scan $::extMsg nu len] != 1 } { exit } set ::extMsg {} if { [coread $len] != $len } { exit } # Read full message. } } proc Listen {} { # Listening on stdin set ::forever 1 coroutine coro_stdin ExtRead vwait forever } set extMsg {} Listen
问题根源
问题出在Tcl的默认通道行为和长度统计逻辑上:
- Tcl默认将stdin视为文本通道,读取时会自动把字节流按系统默认编码(如UTF-8)解码为Unicode字符。这导致包含多字节UTF-8字符时,
[string length $data]返回的是字符数,而非实际读取的字节数。 - 浏览器Native Messaging前缀的长度是消息UTF-8编码后的总字节数,但你的代码用字符数去计算剩余读取字节数,导致
remBytes被错误减少,最终coread会等待读取比实际需要更多的字节,从而卡住。
修复方案
- 将stdin设置为二进制模式:关闭自动文本解码,让Tcl直接处理原始字节流,这是核心修复点。
- 确保长度统计基于字节数:二进制模式下,
read返回原始字节串,此时string length等价于字节数,也可以用binary scan显式统计字节数。
修改后的关键代码
在Listen过程开头添加二进制模式配置:
proc Listen {} { # 配置stdin为二进制模式,禁止自动解码字节流 chan configure stdin -translation binary set ::forever 1 coroutine coro_stdin ExtRead vwait forever }
若要更严谨地统计字节数,可修改coread中的长度计算逻辑:
set data [read stdin $remBytes] # 显式扫描字节数 binary scan $data c* byteList set byteCount [llength $byteList] append ::extMsg $data set remBytes [expr {$remBytes - $byteCount}]
额外说明
- C语言的
fread直接操作字节,不受编码影响,所以能正常工作;而Tcl的文本模式会合并多字节字符,导致长度统计错误。 - 二进制模式下,
binary scan $::extMsg nu len的nu格式(原生字节序无符号32位)与C语言uint32_t的行为完全一致,可正确解析前缀长度。
内容的提问来源于stack exchange,提问作者Gary
相关产品推荐
相关产品推荐

