如何从加拿大环境部气象站点获取露点与风速数据?
问题与解决方案
问题描述
无法从目标天气页面获取露点(Dew Point)和风速(Wind)数据,尝试了基于requests和lxml库的Python代码但未成功。原尝试代码如下:
# Get the html page resp=requests.get("https://weather.gc.ca/city/pages/ab-52_metric_e.html") # Build html tree html_tree=html.fromstring(resp.text) #Dew_point=html_tree.xpath("//dd[@class='mrgn-bttm-0 wxo-metric-hide'][(parent::dl[@class='dl-horizontal wxo-conds-col2'])]//text()")[1].replace("Â", "") # Print Dew_point #print(f"Dew_point in {city_name} is {Dew_point}") #Wind=html_tree.xpath("//dd[@class='longContent mrgn-bttm-0 wxo-metric-hide'][(parent::dl[@class='dl-horizontal wxo-conds-col2'])]//text()")[0].replace("Â", "") # Print Wind #print(f"Wind in {city_name} is {Wind}")
期望获取的数据格式:
Dew point:-2.3°C Wind: NE 9 km/h
对应的目标HTML片段:
<dt>Temperature:</dt> <dd class="mrgn-bttm-0 wxo-metric-hide">13.2°<abbr title="Celsius">C</abbr> </dd> <dd class="mrgn-bttm-0 wxo-imperial-hide wxo-city-hidden">55.8° <abbr title="Fahrenheit">F</abbr> </dd> <dt>Dew point:</dt> <dd class="mrgn-bttm-0 wxo-metric-hide">-2.3°<abbr title="Celsius">C</abbr> </dd> <dd class="mrgn-bttm-0 wxo-imperial-hide wxo-city-hidden">27.9°<abbr title="Fahrenheit">F</abbr> </dd> <dt>Humidity:</dt> <dd class="mrgn-bttm-0">34%</dd> </dl></div> <div class="col-sm-4"><dl class="dl-horizontal wxo-conds-col3"> <dt>Wind:</dt> <dd class="longContent mrgn-bttm-0 wxo-metric-hide"> <abbr title="Northeast">NE</abbr> 9 <abbr title="kilometres per hour">km/h</abbr> </dd> <dd class="longContent mrgn-bttm-0 wxo-imperial-hide wxo-city-hidden"> <abbr title="Northeast">NE</abbr> 6 <abbr title="miles per hour">mph</abbr> </dd> <dt>Visibility:</dt> <dd class="mrgn-bttm-0 wxo-metric-hide">48 <abbr title="kilometres">km</abbr> </dd> <dd class="mrgn-bttm-0 wxo-imperial-hide wxo-city-hidden">30 miles</dd> </dl></div>
解决方案
原代码问题分析
- 风速对应的
<dl>元素class是wxo-conds-col3而非wxo-conds-col2,原XPath父元素定位错误。 - 通过索引取文本的方式不稳定,页面结构变化会导致取值错误。
- 未处理页面编码,导致出现
Â这类乱码字符。
修正后的代码
import requests from lxml import html # 请求页面并指定编码 resp = requests.get("https://weather.gc.ca/city/pages/ab-52_metric_e.html") resp.encoding = 'utf-8' # 显式指定编码避免乱码 html_tree = html.fromstring(resp.text) # 获取露点数据:定位<dt>文本为"Dew point:"的下一个<dd>(公制) dew_point_ele = html_tree.xpath("//dt[text()='Dew point:']/following-sibling::dd[@class='mrgn-bttm-0 wxo-metric-hide']")[0] dew_point = ''.join(dew_point_ele.xpath(".//text()")).strip() # 获取风速数据:定位<dt>文本为"Wind:"的下一个<dd>(公制) wind_ele = html_tree.xpath("//dt[text()='Wind:']/following-sibling::dd[@class='longContent mrgn-bttm-0 wxo-metric-hide']")[0] wind = ''.join(wind_ele.xpath(".//text()")).strip() # 输出结果 print(f"Dew point:{dew_point}") print(f"Wind: {wind}")
代码说明
- 通过
<dt>的文本内容精准定位对应的<dd>元素,避免依赖父元素class或索引的不稳定问题。 - 使用
''.join(ele.xpath(".//text()"))提取元素内所有文本并拼接,保留完整的单位信息。 - 显式设置
resp.encoding = 'utf-8'解决乱码问题,无需手动替换特殊字符。
内容的提问来源于stack exchange,提问作者laisp-lostlamb
相关产品推荐
相关产品推荐

