如何用BeautifulSoup从复杂HTML代码中提取温度数值?
提取HTML表格中的温度数值
解决思路
定位到目标<tr>元素后,跳过首个标题列的<td>,遍历剩余单元格,通过正则表达式提取其中的正负整数(或浮点数)温度值。同时可修复请求编码问题,避免出现Â这类乱码字符。
修改后的代码示例
import re import requests from bs4 import BeautifulSoup def get_avg_high_temp(*args): url_site = get_url(*args) # 假设get_url已正确返回目标URL for url in url_site: res = requests.get(url) # 修复编码问题,避免出现Â乱码 res.encoding = 'utf-8' soup = BeautifulSoup(res.text, "html.parser") temp_row = soup.find(title='Temp.') if not temp_row: print("未找到温度行") continue # 跳过第一个包含"Temp."的td,获取所有温度单元格 temp_cells = temp_row.find_all('td')[1:] temperatures = [] for cell in temp_cells: cell_text = cell.get_text(strip=True) # 匹配正负整数,若需支持小数可改用正则:r'-?\d+\.?\d*' match = re.search(r'-?\d+', cell_text) if match: # 转换为整数,如需浮点则用float() temp = int(match.group()) temperatures.append(temp) print("提取的温度数值:", temperatures) get_avg_high_temp('Calgary Canada')
代码说明
res.encoding = 'utf-8':强制设置响应编码为UTF-8,解决网页字符编码不匹配导致的Â乱码问题,让温度文本显示为正常的-6 °C。temp_row.find_all('td')[1:]:跳过第一个<td>(标题列),只处理包含温度的单元格。- 正则表达式
r'-?\d+':匹配包含负号的整数,若网页中存在小数温度,可替换为r'-?\d+\.?\d*'来匹配浮点数。 - 提取到的字符串通过
int()或float()转换为数值类型,方便后续计算或使用。
内容的提问来源于stack exchange,提问作者WQureshi
相关产品推荐
相关产品推荐

