You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup筛选网页表格中指定类型的href属性链接

你可以通过匹配a标签所属表格行内的图片标识实现过滤,自动站对应行的图片src会包含automatica相关关键词,手动站对应为manual相关,筛选符合条件的链接即可。

修改后的完整代码如下:

import urllib2
import re
from bs4 import BeautifulSoup

# Fetch URL
url = 'http://meteo.navarra.es/estaciones/descargardatos.cfm'

request = urllib2.Request(url)
request.add_header('Accept-Encoding', 'utf-8')

# Response has UTF-8 charset header, and HTML body which is UTF-8 encoded
response = urllib2.urlopen(request)

# Parse with BeautifulSoup
soup = BeautifulSoup(response,'html.parser')

for a in soup.find_all('a',{'href': re.compile(r'descargardatos_estacion.*')}):
    # 获取当前a标签所属的表格行
    parent_tr = a.find_parent('tr')
    # 查找该行内的图片元素,检查是否为自动站标识
    img_tag = parent_tr.find('img', src=re.compile(r'automatic', re.IGNORECASE))
    if img_tag:
        estacion = 'http://meteo.navarra.es/estaciones/' + a.attrs.get('href')
        print(estacion)
        # descarga_csvs(estacion)

如果实际页面中自动站图片的src关键词和示例中的automatic不一致,你可以把正则里的关键词替换成实际匹配的内容即可。

内容的提问来源于stack exchange,提问作者aarribas12

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 18:36:07