You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup从复杂HTML代码中提取温度数值?

提取HTML表格中的温度数值

解决思路

定位到目标<tr>元素后,跳过首个标题列的<td>,遍历剩余单元格,通过正则表达式提取其中的正负整数(或浮点数)温度值。同时可修复请求编码问题,避免出现Â这类乱码字符。

修改后的代码示例

import re
import requests
from bs4 import BeautifulSoup

def get_avg_high_temp(*args):
    url_site = get_url(*args)  # 假设get_url已正确返回目标URL
    for url in url_site:
        res = requests.get(url)
        # 修复编码问题,避免出现Â乱码
        res.encoding = 'utf-8'
        soup = BeautifulSoup(res.text, "html.parser")
        temp_row = soup.find(title='Temp.')
        if not temp_row:
            print("未找到温度行")
            continue
        
        # 跳过第一个包含"Temp."的td,获取所有温度单元格
        temp_cells = temp_row.find_all('td')[1:]
        temperatures = []
        
        for cell in temp_cells:
            cell_text = cell.get_text(strip=True)
            # 匹配正负整数,若需支持小数可改用正则:r'-?\d+\.?\d*'
            match = re.search(r'-?\d+', cell_text)
            if match:
                # 转换为整数,如需浮点则用float()
                temp = int(match.group())
                temperatures.append(temp)
        
        print("提取的温度数值:", temperatures)

get_avg_high_temp('Calgary Canada')

代码说明

  • res.encoding = 'utf-8':强制设置响应编码为UTF-8,解决网页字符编码不匹配导致的Â乱码问题,让温度文本显示为正常的-6 °C。
  • temp_row.find_all('td')[1:]:跳过第一个<td>(标题列),只处理包含温度的单元格。
  • 正则表达式r'-?\d+':匹配包含负号的整数,若网页中存在小数温度,可替换为r'-?\d+\.?\d*'来匹配浮点数。
  • 提取到的字符串通过int()或float()转换为数值类型,方便后续计算或使用。

内容的提问来源于stack exchange,提问作者WQureshi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 06:35:23