如何用Bash脚本提取HTML表格指定数据生成CSV?
Bash解析HTML表格的问题解决
问题说明
已成功获取目标HTML表格,但数据提取效果不符合预期,需要生成指定格式的CSV内容,每条记录格式为地区;邮编;坐标。
目标HTML表格片段
<tr><td><small>1</small></td><td>Kalisz</td><td>62-800</td><td>Poland</td><td>Greater Poland</td><td>Kalisz</td><td>Kalisz<tr><td></td><td colspan=6> <a href="/maps/browse_51.75_18.087.html" rel="nofollow"><small>51.75/18.087</small></a></td></tr> <tr class="odd"><td><small>2</small></td><td>Piotrków Trybunalski</td><td>97-300</td><td>Poland</td><td>Łódź Voivodeship</td><td>Piotrków Trybunalski</td><td>Piotrków Trybunalski<tr class="odd"><td></td><td colspan=6> <a href="/maps/browse_51.411_19.689.html" rel="nofollow"><small>51.411/19.689</small></a></td></tr> <tr><td><small>3</small></td><td>Toruń</td><td>87-100</td><td>Poland</td><td>Kujawsko-Pomorskie</td><td>Toruń</td><td>Toruń<tr><td></td><td colspan=6> <a href="/maps/browse_53.021_18.623.html" rel="nofollow"><small>53.021/18.623</small></a></td></tr>
现有Bash脚本
#!/bin/bash URL="https://www.geonames.org/postalcode-search.html?country=PL&q=" HTML=$(curl -s "$URL") (echo "$HTML" | grep -A 201 "<table class=\"restable\">" | tail -n 200 )>> table.html html_lines=() while IFS= read -r line; do html_lines+=("$line") done < "table.html" for html_line in "${html_lines[@]}"; do field1_value=$(echo "$html_line" | grep -oP '(?<=<td>)[^<]+(?=</td>)') field2_value=$(echo "$html_line" | grep -oP '[0-9]{2}-[0-9]{3}') field3_value=$(echo "$html_line" | grep -oP '(?<=<small>)[^<]+(?=</small>)') # Printing the extracted fields echo "$field1_value;$field2_value;$field3_value" >> output.txt done
当前输出
Kalisz 62-800 Poland Greater Poland Kalisz;62-800;1 51.75/18.087 Piotrków Trybunalski 97-300 Poland Łódź Voivodeship Piotrków Trybunalski;97-300;2 51.411/19.689 Toruń 87-100 Poland Kujawsko-Pomorskie Toruń;87-100;3 53.021/18.623
期望输出
Greater Poland;62-800;51.75/18.087 Łódź Voivodeship;97-300;51.411/19.689 Kujawsko-Pomorskie;87-100;53.021/18.623
解决方案
原脚本核心问题是逐行处理导致数据拆分,无法将主数据行的地区、邮编与下一行的坐标关联,且字段匹配范围过广。以下是改进后的脚本:
#!/bin/bash URL="https://www.geonames.org/postalcode-search.html?country=PL&q=" # 获取页面内容,过滤表格相关行并提取目标数据 curl -s "$URL" | grep -A 400 '<table class="restable">' | awk ' BEGIN { RS="</tr>"; FS="</td>" } # 处理主数据行(不含colspan的<tr>) /^<tr/ && !/colspan/ { zipcode = "" region_name = "" for (i=1; i<=NF; i++) { # 提取第3个<td>中的邮编 if (i == 3 && match($i, /[0-9]{2}-[0-9]{3}/, zip)) { zipcode = zip[0] } # 提取第5个<td>中的地区名称 if (i == 5 && match($i, />([^<]+)$/, region)) { region_name = region[1] } } } # 处理坐标行(含colspan的<tr>) /colspan/ { if (match($0, /<small>([^<]+)<\/small>/, coord)) { coordinates = coord[1] # 输出组合后的CSV行 print region_name ";" zipcode ";" coordinates } }' > output.txt
脚本说明
- 数据抓取与过滤:用
curl获取页面内容,通过grep定位到目标表格并获取足够多的后续行(适配200条数据)。 - AWK处理逻辑:
- 设置记录分隔符为
</tr>,将每个<tr>块作为独立记录处理。 - 识别主数据行(不含
colspan),精准提取第3个<td>的邮编和第5个<td>的地区名称。 - 识别坐标行(含
colspan),提取<small>标签内的坐标值,然后将地区、邮编、坐标组合成指定格式输出。
- 设置记录分隔符为
内容的提问来源于stack exchange,提问作者Maqcel
相关产品推荐
相关产品推荐

