如何抓取网页指定标签页内的年度降雪嵌入式表格数据?
问题:无法抓取OnTheSnow网站的年度降雪数据
我尝试从OnTheSnow的滑雪场历史降雪页面抓取数据,当前脚本只能获取月度总降雪量,但网页上需要切换到“Annual”标签页才能看到年度总降雪数据,我能成功抓取月度数据,却拿不到年度数据。
以下是我的实现代码:
def historic_snowfall(): #Resort naming convetion on OpenSnow.com # Revelstoke, BC: british-columbia, revelstoke-mountain # Whistler, BC: british-columbia, whistler-blackcomb # Lake Louise, AB: alberta, lake-louise # Big Sky, MT: montana, big-sky-resort # Snowbird, UT: utah, snowbird # Palisades: califronia, squaw-valley-usa # Steamboat, CO: colorado, steamboat # Copper Mountain, CO: colorado, copper-mountain # Aspen, CO: colorado, aspen-snowmass # Jackson Hole, WY: wyoming, jackson-hole # Taos, NM: taos-ski-valley resort_table = ['alberta/lake-louise', 'montana/big-sky-resort', 'utah/snowbird', 'california/squaw-valley-usa', 'colorado/steamboat', 'colorado/copper-mountain', 'colorado/aspen-snowmass', 'wyoming/jackson-hole', 'new-mexico/taos-ski-valley' ] resort = (resort_table[1]) hsnow_url = 'https://www.onthesnow.com/' + resort + '/historical-snowfall' page = requests.get(hsnow_url) soup = BeautifulSoup(page.content, "html.parser") data = [] table = soup.find('table') # Find table table_body = table.find('tbody') # Find body of table rows = table_body.find_all('tr') #find rows within table for row in rows: cols = row.find_all('td') # within rows pull column data cols = [ele.text.strip() for ele in cols] data.append([ele for ele in cols if ele]) # Get rid of empty values return data #returns monthly averages not yearly results
解决方案
问题原因
网页的月度、年度降雪数据其实已经全部加载在页面中,只是年度数据所在的标签页默认被隐藏,你的脚本默认抓取了第一个显示的表格(即月度数据)。
修改后的代码
import requests from bs4 import BeautifulSoup def historic_snowfall(): #Resort naming convetion on OpenSnow.com # Revelstoke, BC: british-columbia, revelstoke-mountain # Whistler, BC: british-columbia, whistler-blackcomb # Lake Louise, AB: alberta, lake-louise # Big Sky, MT: montana, big-sky-resort # Snowbird, UT: utah, snowbird # Palisades: califronia, squaw-valley-usa # Steamboat, CO: colorado, steamboat # Copper Mountain, CO: colorado, copper-mountain # Aspen, CO: colorado, aspen-snowmass # Jackson Hole, WY: wyoming, jackson-hole # Taos, NM: taos-ski-valley resort_table = ['alberta/lake-louise', 'montana/big-sky-resort', 'utah/snowbird', 'california/squaw-valley-usa', 'colorado/steamboat', 'colorado/copper-mountain', 'colorado/aspen-snowmass', 'wyoming/jackson-hole', 'new-mexico/taos-ski-valley' ] resort = resort_table[1] hsnow_url = f'https://www.onthesnow.com/{resort}/historical-snowfall' page = requests.get(hsnow_url) soup = BeautifulSoup(page.content, "html.parser") data = [] # 定位年度数据所在的标签容器(优先通过id定位) annual_container = soup.find('div', id='annual') # 若id不匹配,尝试通过标签关联属性定位 if not annual_container: annual_container = soup.find('div', class_='tab-pane', attrs={'aria-labelledby': 'annual-tab'}) if annual_container: # 从容器内提取表格数据 table = annual_container.find('table') table_body = table.find('tbody') rows = table_body.find_all('tr') for row in rows: cols = row.find_all('td') cols = [ele.text.strip() for ele in cols] data.append([ele for ele in cols if ele]) else: print("未找到年度数据表格,请检查页面结构") return data # 返回年度降雪数据
额外说明
如果上述选择器无法定位到数据,建议手动检查目标页面的HTML结构:
- 打开目标页面并切换到“Annual”标签页
- 右键选择「检查」,找到表格元素,查看其所在父容器的id、class等属性
- 调整代码中的选择器以匹配实际结构
如果页面采用动态加载(即年度数据需点击后才从服务器获取),则需要使用selenium等工具模拟浏览器点击行为后再抓取数据。
内容的提问来源于stack exchange,提问作者chefboyrb
相关产品推荐
相关产品推荐

