You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取含隐藏类的网页表格全部数据?(附Python爬取代码)

解决表格嵌套span内容缺失的问题

问题原因

Pandas的read_html工具默认只会提取单元格的直接文本内容,不会解析嵌套在<span>这类标签里的内容,所以你看到的class为ntee的span(对应数字4)会被忽略。

方案一:手动遍历表格提取完整内容

用BeautifulSoup先定位表格,逐个单元格处理嵌套标签,再生成DataFrame,这种方式最可控:

import requests
from bs4 import BeautifulSoup
import pandas as pd

# 替换成你的目标网页URL
target_url = "https://your-target-weather-page.com"
resp = requests.get(target_url)
soup = BeautifulSoup(resp.text, "html.parser")

# 定位目标表格(根据实际页面调整选择器,比如用id、class或位置)
weather_table = soup.find("table")  # 示例:取页面第一个表格,可改成find('table', class_='weather')

# 处理每一行和单元格
processed_data = []
for row in weather_table.find_all("tr"):
    cell_contents = []
    for cell in row.find_all(["td", "th"]):
        # 优先提取ntee类span的内容
        ntee_tag = cell.find("span", class_="ntee")
        if ntee_tag:
            cell_contents.append(ntee_tag.get_text(strip=True))
        else:
            # 没有嵌套标签就取单元格本身的文本
            cell_contents.append(cell.get_text(strip=True))
    processed_data.append(cell_contents)

# 生成DataFrame(第一行做表头,后面是数据)
weather_df = pd.DataFrame(processed_data[1:], columns=processed_data[0])
print(weather_df)

方案二:先修改HTML再用read_html读取

如果习惯用read_html,可以先把页面里的ntee标签替换成纯文本,再让Pandas解析:

import requests
from bs4 import BeautifulSoup
import pandas as pd

target_url = "https://your-target-weather-page.com"
resp = requests.get(target_url)
soup = BeautifulSoup(resp.text, "html.parser")

# 遍历所有ntee标签,替换成标签内的文本
for span in soup.find_all("span", class_="ntee"):
    span.replace_with(span.get_text(strip=True))

# 用read_html读取处理后的页面
dfs = pd.read_html(str(soup))
# 取目标表格(比如第一个,根据实际情况调整索引)
weather_df = dfs[0]
print(weather_df)

注意事项

  • 调整表格选择器:如果页面有多个表格,要给find("table")加更精准的条件(比如id="weather-table"或class="forecast-table"),避免选错表格。
  • 测试文本提取:如果还有其他嵌套标签的内容需要提取,可以在单元格处理逻辑里添加对应判断。

内容的提问来源于stack exchange,提问作者Stetco Oana

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 17:24:47