You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Bash脚本提取HTML表格指定数据生成CSV?

Bash解析HTML表格的问题解决

问题说明

已成功获取目标HTML表格,但数据提取效果不符合预期,需要生成指定格式的CSV内容,每条记录格式为地区;邮编;坐标。

目标HTML表格片段

<tr><td><small>1</small></td><td>Kalisz</td><td>62-800</td><td>Poland</td><td>Greater Poland</td><td>Kalisz</td><td>Kalisz<tr><td></td><td colspan=6>&nbsp;&nbsp;&nbsp;<a href="/maps/browse_51.75_18.087.html" rel="nofollow"><small>51.75/18.087</small></a></td></tr>
<tr class="odd"><td><small>2</small></td><td>Piotrków Trybunalski</td><td>97-300</td><td>Poland</td><td>Łódź Voivodeship</td><td>Piotrków Trybunalski</td><td>Piotrków Trybunalski<tr class="odd"><td></td><td colspan=6>&nbsp;&nbsp;&nbsp;<a href="/maps/browse_51.411_19.689.html" rel="nofollow"><small>51.411/19.689</small></a></td></tr>
<tr><td><small>3</small></td><td>Toruń</td><td>87-100</td><td>Poland</td><td>Kujawsko-Pomorskie</td><td>Toruń</td><td>Toruń<tr><td></td><td colspan=6>&nbsp;&nbsp;&nbsp;<a href="/maps/browse_53.021_18.623.html" rel="nofollow"><small>53.021/18.623</small></a></td></tr>

现有Bash脚本

#!/bin/bash

URL="https://www.geonames.org/postalcode-search.html?country=PL&q="

HTML=$(curl -s "$URL")

(echo "$HTML" | grep -A 201 "<table class=\"restable\">" | tail -n 200 )>> table.html

html_lines=()
while IFS= read -r line; do
  html_lines+=("$line")
done < "table.html"

for html_line in "${html_lines[@]}"; do
  field1_value=$(echo "$html_line" | grep -oP '(?<=<td>)[^<]+(?=</td>)')
  field2_value=$(echo "$html_line" | grep -oP '[0-9]{2}-[0-9]{3}')
  field3_value=$(echo "$html_line" | grep -oP '(?<=<small>)[^<]+(?=</small>)')

  # Printing the extracted fields
  echo "$field1_value;$field2_value;$field3_value" >> output.txt
done

当前输出

Kalisz
62-800
Poland
Greater Poland
Kalisz;62-800;1
51.75/18.087
Piotrków Trybunalski
97-300
Poland
Łódź Voivodeship
Piotrków Trybunalski;97-300;2
51.411/19.689
Toruń
87-100
Poland
Kujawsko-Pomorskie
Toruń;87-100;3
53.021/18.623

期望输出

Greater Poland;62-800;51.75/18.087
Łódź Voivodeship;97-300;51.411/19.689
Kujawsko-Pomorskie;87-100;53.021/18.623

解决方案

原脚本核心问题是逐行处理导致数据拆分,无法将主数据行的地区、邮编与下一行的坐标关联,且字段匹配范围过广。以下是改进后的脚本:

#!/bin/bash

URL="https://www.geonames.org/postalcode-search.html?country=PL&q="

# 获取页面内容,过滤表格相关行并提取目标数据
curl -s "$URL" | grep -A 400 '<table class="restable">' | awk '
BEGIN { RS="</tr>"; FS="</td>" }
# 处理主数据行(不含colspan的<tr>)
/^<tr/ && !/colspan/ {
    zipcode = ""
    region_name = ""
    for (i=1; i<=NF; i++) {
        # 提取第3个<td>中的邮编
        if (i == 3 && match($i, /[0-9]{2}-[0-9]{3}/, zip)) {
            zipcode = zip[0]
        }
        # 提取第5个<td>中的地区名称
        if (i == 5 && match($i, />([^<]+)$/, region)) {
            region_name = region[1]
        }
    }
}
# 处理坐标行(含colspan的<tr>)
/colspan/ {
    if (match($0, /<small>([^<]+)<\/small>/, coord)) {
        coordinates = coord[1]
        # 输出组合后的CSV行
        print region_name ";" zipcode ";" coordinates
    }
}' > output.txt

脚本说明

  1. 数据抓取与过滤:用curl获取页面内容,通过grep定位到目标表格并获取足够多的后续行(适配200条数据)。
  2. AWK处理逻辑:
    • 设置记录分隔符为</tr>,将每个<tr>块作为独立记录处理。
    • 识别主数据行(不含colspan),精准提取第3个<td>的邮编和第5个<td>的地区名称。
    • 识别坐标行(含colspan),提取<small>标签内的坐标值,然后将地区、邮编、坐标组合成指定格式输出。

内容的提问来源于stack exchange,提问作者Maqcel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 18:15:55