使用BS4(Python)爬取天气数据时DataFrame返回空值求助
解决从meteo.gov.lk提取天气数据为空DataFrame的问题
问题情况
尝试从指定页面提取天气数据并导出至DataFrame,但运行代码后输出为空DataFrame。
原代码:
import pandas as pd import requests from bs4 import BeautifulSoup url = "http://meteo.gov.lk/index.php?option=com_content&view=article&id=102&Itemid=360&lang=en" data = requests.get(url).text soup = BeautifulSoup(data, 'html5lib') df = pd.DataFrame(columns=["Location", "Status", "Temperature", "Rainfall", "Reported_Time"]) for row in soup.find_all('tr'): col = row.find_all("td") Location = col[0].text Status = col[1].text Temperature = col[2].text Rainfall = col[3].text Reported_Time = col[4].text df = df.append({"Location":Location,"Status":Status,"Temperature":Temperature,"Rainfall":Rainfall,"Reported_Time":Reported_Time,}, ignore_index=True) print(df)
运行输出:
Empty DataFrame Columns: [Location, Status, Temperature, Rainfall,RH, Reported_Time] Index: []
问题原因
- 页面存在大量无关的
<tr>标签,直接遍历所有<tr>无法定位到天气数据行 - 天气数据表格需精准定位,且表头行用
<th>标签而非<td>,直接处理会导致空数据或索引错误 df.append()已被Pandas弃用,且效率低下
修复后的代码
import pandas as pd import requests from bs4 import BeautifulSoup url = "http://meteo.gov.lk/index.php?option=com_content&view=article&id=102&Itemid=360&lang=en" response = requests.get(url) response.encoding = 'utf-8' # 确保编码正确避免乱码 soup = BeautifulSoup(response.text, 'html5lib') # 精准定位天气数据所在表格(匹配页面实际表格类名) weather_table = soup.find('table', attrs={'class': 'table table-bordered'}) if not weather_table: print("未找到天气数据表格") exit() # 用列表存储数据,后续一次性转DataFrame更高效 data_rows = [] # 跳过表头行,只处理数据行 for row in weather_table.find_all('tr')[1:]: cols = row.find_all('td') # 确保当前行有足够列数,避免索引报错 if len(cols) >= 5: location = cols[0].text.strip() status = cols[1].text.strip() temperature = cols[2].text.strip() rainfall = cols[3].text.strip() reported_time = cols[4].text.strip() data_rows.append({ "Location": location, "Status": status, "Temperature": temperature, "Rainfall": rainfall, "Reported_Time": reported_time }) df = pd.DataFrame(data_rows, columns=["Location", "Status", "Temperature", "Rainfall", "Reported_Time"]) print(df)
关键修改说明
- 精准定位表格:通过类名匹配找到天气数据专属表格,避免遍历无关的页面元素
- 跳过表头:从第2行开始遍历表格,跳过用
<th>标签的表头行 - 编码处理:设置响应编码为UTF-8,防止特殊字符乱码
- 异常防护:判断列数是否足够,避免部分数据行缺失导致索引错误
- 高效构建DataFrame:用列表暂存数据后一次性转换,替代已弃用的
df.append()方法
内容的提问来源于stack exchange,提问作者Snyder Fox
相关产品推荐
相关产品推荐

