Bash脚本需求:统计指定字符串中word前的字符数并格式化输出
解决Bash脚本统计"word"前置字符数的问题
Hey there! Let's get your script working exactly how you need it. First, let's fix the tiny syntax bug in your initial code: when assigning the string to a variable, you don't use the $ prefix. So that line should be str="word blah-blah word one more time word again" instead of $word=....
Now, here are two solid approaches to get the output you're after:
方法1:使用awk处理(简洁高效)
awk是处理文本分割和统计的绝佳工具,我们可以用它按"word"分割字符串,然后逐个计算每个"word"前的字符长度:
#!/bin/bash str="word blah-blah word one more time word again" awk -F'word' '{ # 遍历分割后的所有字段,从第2个开始对应每个"word" for (i=2; i<=NF; i++) { # 前一个字段就是当前"word"之前的内容 prev_content = $(i-1) prev_length = length(prev_content) printf "word %d ##%s\n", prev_length, prev_content } }' <<< "$str"
代码解释:
-F'word':把输入字符串以"word"作为分隔符拆分- 循环从
i=2开始:因为第一个字段是第一个"word"之前的空内容,对应第一个输出行 length(prev_content):计算前置内容的字符数printf:严格按照你要求的格式输出结果
方法2:纯Bash正则匹配(无需额外工具)
如果不想依赖awk,纯Bash的正则匹配也能实现这个需求:
#!/bin/bash str="word blah-blah word one more time word again" # 循环匹配每个"word"及其前置内容 while [[ "$str" =~ (.*?)word(.*) ]]; do # 捕获组1是"word"之前的内容,组2是剩余字符串 prev_content="${BASH_REMATCH[1]}" prev_length="${#prev_content}" printf "word %d ##%s\n" "$prev_length" "$prev_content" # 更新字符串为"word"之后的部分,继续循环 str="${BASH_REMATCH[2]}" done
代码解释:
(.*?)word(.*):非贪婪正则匹配,每次找到最靠前的"word",捕获它前面的内容和后面的剩余字符串${#prev_content}:Bash中获取字符串长度的语法- 循环直到字符串中没有"word"为止
输出结果
两种方法都会输出你需要的格式:
word 0 ## word 9 ##blah-blah word 11 ##one more time
内容的提问来源于stack exchange,提问作者Nick
相关产品推荐
相关产品推荐

