如何用Shell Scripting计算文件行平均值并按指定格式输出?
Alright, let's walk through exactly how to build a shell script that handles this docx processing task. Here's a step-by-step solution tailored to your needs:
Step 1: Convert .docx to Plain Text
Shell can't directly read .docx files, so first we need to convert it to a plain text format. We'll use pandoc (a versatile document converter) for this—if you don't have it installed, grab it via your package manager:
- Debian/Ubuntu:
sudo apt install pandoc - CentOS/RHEL:
sudo dnf install pandoc
Run this command to convert the docx to text:
pandoc -s file1.docx -o file1.txt
Alternative: If pandoc isn't available, you can use antiword (a tool specifically for extracting text from Word docs):
antiword file1.docx > file1.txt
Step 2: Build the Shell Script
Create a script (e.g., process_doc.sh) with the following content. It handles date conversion, average calculations, and output formatting:
#!/bin/bash # Convert docx to text (uncomment your preferred method) pandoc -s file1.docx -o file1.txt # antiword file1.docx > file1.txt # Extract and reformat the date from Line D # Assumes Line D starts with "行D" and the date is the second field formatted_date=$(grep "^行D" file1.txt | sed 's/-/\//g' | awk '{print $2}') # Calculate average for Line B (handles all numeric fields after the first) avg_b=$(grep "^行B" file1.txt | awk '{ sum = 0 for (i=2; i<=NF; i++) sum += $i printf "%.2f", sum / (NF - 1) }') # Calculate average for Line P (same logic as Line B) avg_p=$(grep "^行P" file1.txt | awk '{ sum = 0 for (i=2; i<=NF; i++) sum += $i printf "%.2f", sum / (NF - 1) }') # Generate the final formatted data module cat << EOF > processed_data.txt [数据模块] 日期: $formatted_date B行平均值: $avg_b P行平均值: $avg_p EOF echo "Done! Processed data saved to processed_data.txt"
Step 3: Run the Script
Make the script executable and run it:
chmod +x process_doc.sh ./process_doc.sh
Key Notes to Adjust for Your Exact Doc Format
- If your lines don't start with "行D"/"行B"/"行P" (e.g., there's leading whitespace), remove the
^in thegrepcommands (usegrep "行D"instead ofgrep "^行D"). - If the numeric values/date are in different fields, tweak the
awkfield indices (e.g., if the date is the 3rd field, change$2to$3). - The
printf "%.2f"ensures averages are rounded to 2 decimal places, matching your example (15.63, 29.87).
内容的提问来源于stack exchange,提问作者kittensfurdays

