使用BeautifulSoup从HTML页面提取指定天气预报链接的实现问题
正确实现代码
import requests from bs4 import BeautifulSoup # 请求目标页面 request_url = "https://forecast.weather.gov/MapClick.php?w0=t&w1=td&w2=wc&w3=sfcwind&w3u=1&w4=sky&w5=pop&w6=rh&w7=rain&w8=thunder&w9=snow&w10=fzg&w11=sleet&AheadHour=3&Submit=Submit&FcstType=graphical&textField1=40.7328&textField2=-73.2177&site=okx&unit=0&dd=&bw=" page = requests.get(request_url, timeout=10) soup = BeautifulSoup(page.content, 'html.parser') target_links = [] # 遍历所有带href属性的a标签 for a_tag in soup.find_all("a", href=True): # 目标链接核心特征是包含FcstType=digital参数,直接过滤即可 if "FcstType=digital" in a_tag["href"]: # 自动补全相对路径为完整绝对链接 full_link = requests.compat.urljoin(request_url, a_tag["href"]) target_links.append(full_link) # 对重复链接去重后输出 target_links = list(set(target_links)) for link in target_links: print(link)
原代码错误说明
- 查找带指定alt属性的元素语法错误,
soup.find(alt="xxx")不是合法写法,需要通过attrs参数指定属性过滤,正确写法为soup.find(attrs={"alt": "Hourly weather graph of forecast elements. Click for text representation."}) - 遍历逻辑错误,循环内反复调用
soup.find('area')只会返回第一个匹配的area标签,没有实际遍历所有a标签 print(x)["href"]语法错误,print函数返回None,无法从中获取href属性- 没有对链接做特征过滤,不需要限制标签位置,直接通过目标链接独有的参数特征匹配即可
内容的提问来源于stack exchange,提问作者alexsey
相关产品推荐
相关产品推荐

