请求编写Bash脚本统计报纸文本字符并生成CSV
Bash脚本开发:统计报纸文章字符数
需求概述
处理一个8万行、5.6MB的UTF-8编码法语文本文件,实现以下目标:
- 文件包含约900篇报纸文章,每篇以
news01、news02这类报社标识行开头 - 忽略标识行后对应报社的固定数量标题行,以及文章末尾固定数量的免责声明行
- 统计标识行之后、免责声明行之前的文本字符数(含字母、数字、标点、空格,不含
\n) - 处理文本时需将每行首尾的空格/制表符缩减至最多一个空格
- 输出格式为CSV,每行是
报社标识;总字符数
报社配置规则
用Bash关联数组存储各报社的参数,示例如下:
declare -A INFOS INFOS[news01N]="Newspaper 1 name" # 报社名称(可选,本需求未用到) INFOS[news01H]=1 # 需要忽略的标题行数 INFOS[news01D]=2 # 需要忽略的免责声明行数 INFOS[news02N]="Newspaper 2 name" INFOS[news02H]=2 INFOS[news02D]=1
样本输入
news01 this line to be ignored too this is the first article of the file disclaimer to be ignored news02 these lines to be ignored too this is the second article of the file disclaimer to be ignored
样本输出
news01;37 news02;38
用户现有代码
#! /bin/bash while IFS= read -r line; do if [[ "$line" != *"news01"* ]]; then echo ${#line}; fi done < $1
完整解决方案脚本
#!/bin/bash # 配置各报社的标题跳过行数(H)和免责声明跳过行数(D) declare -A INFOS INFOS[news01H]=1 INFOS[news01D]=2 INFOS[news02H]=2 INFOS[news02D]=1 # 初始化状态变量 current_news="" skip_head=0 remaining_disclaimer=0 article_lines=() # 逐行读取输入文件 while IFS= read -r line; do # 检测是否为新文章的报社标识行 if [[ "$line" =~ ^news[0-9]+$ ]]; then # 输出上一篇文章的统计结果(如果存在) if [[ -n "$current_news" ]]; then # 将所有有效行用单个空格连接,计算总字符数 full_article=$(IFS=' '; echo "${article_lines[*]}") total_chars=${#full_article} echo "${current_news};${total_chars}" fi # 重置状态,开始处理新文章 current_news="$line" skip_head=${INFOS[${current_news}H]} remaining_disclaimer=${INFOS[${current_news}D]} article_lines=() continue fi # 跳过标题行 if [[ $skip_head -gt 0 ]]; then ((skip_head--)) continue fi # 跳过免责声明行 if [[ $remaining_disclaimer -gt 0 ]]; then ((remaining_disclaimer--)) continue fi # 处理文章内容行:修剪首尾空白,非空行加入数组 trimmed_line=$(echo "$line" | sed -e 's/^[[:space:]]*//' -e 's/[[:space:]]*$//') if [[ -n "$trimmed_line" ]]; then article_lines+=("$trimmed_line") fi done # 处理文件中最后一篇文章的统计 if [[ -n "$current_news" ]]; then full_article=$(IFS=' '; echo "${article_lines[*]}") total_chars=${#full_article} echo "${current_news};${total_chars}" fi
脚本说明
- 配置部分:关联数组
INFOS存储各报社的跳过规则,后续新增报社只需添加对应键值对即可 - 状态跟踪:通过
current_news记录当前处理的报社,skip_head和remaining_disclaimer跟踪剩余需跳过的行数 - 内容处理:对每行做首尾空白修剪,非空行存入数组,最后用单个空格连接所有行,自动替换原换行符为空格,符合统计要求
- 收尾处理:循环结束后单独处理最后一篇文章,避免遗漏
内容的提问来源于stack exchange,提问作者Ivan
相关产品推荐
相关产品推荐

