You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页表格爬取问题:无法获取Cellmapper的Band Number数据

解决Cellmapper表格数据爬取返回NaN的问题

你的代码返回全NaN是因为pd.io.html.read_html没有正确识别目标表格的列结构——页面结果表格的每行包含两个<td>标签(一个是描述项,一个是对应值),直接传入所有表格的字符串会导致解析逻辑偏差。

修正方案1:手动解析表格行

这种方式更灵活,能精准提取所需内容:

import requests
import pandas as pd
from bs4 import BeautifulSoup

def htmltodf(url):
    page = requests.get(url)
    soup = BeautifulSoup(page.text, "lxml")
    # 定位到结果表格(通过class匹配)
    result_table = soup.find('table', class_='table table-striped table-hover table-condensed')
    rows = result_table.find_all('tr')
    
    data = []
    for row in rows:
        cols = row.find_all('td')
        if len(cols) == 2:
            label = cols[0].get_text(strip=True)
            value = cols[1].get_text(strip=True)
            data.append({'描述': label, '值': value})
    
    df = pd.DataFrame(data)
    # 获取Band Number
    band_number = df.loc[df['描述'] == 'Band Number', '值'].values[0]
    print(df)
    print(f"Band Number: {band_number}")

# 测试LTE类型,修改url参数可切换net类型和ARFCN值
htmltodf("https://www.cellmapper.net/arfcn?net=LTE&ARFCN=78&MCC=0")

修正方案2:指定表格class用pandas解析

直接通过read_html的attrs参数定位目标表格,简化代码:

import requests
import pandas as pd

def htmltodf(url):
    page = requests.get(url)
    # 匹配目标表格的class,读取表格数据
    dfs = pd.read_html(page.text, attrs={'class': 'table table-striped table-hover table-condensed'})
    if dfs:
        df = dfs[0]
        df.columns = ['描述', '值']
        print(df)
        band_number = df.loc[df['描述'] == 'Band Number', '值'].values[0]
        print(f"Band Number: {band_number}")

htmltodf("https://www.cellmapper.net/arfcn?net=LTE&ARFCN=78&MCC=0")

切换网络类型(LTE/3G/2G)或修改ARFCN值时,只需调整URL中对应的参数即可。

内容的提问来源于stack exchange,提问作者Alibenc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 21:05:32