You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于BeautifulSoup抓取ISP电缆调制解调器信道上下行数据生成CSV

解决方案:抓取电缆调制解调器信道数据并导出为CSV

看起来你已经迈出了第一步,但原代码只完成了页面请求和解析的初始部分,缺少定位目标表格、提取信道数据、写入CSV文件这几个核心环节。我帮你完善整个流程,下面是可以直接运行的完整代码,同时会解释关键步骤。

import requests
from bs4 import BeautifulSoup
import csv
from datetime import datetime

def fetch_modem_channel_data():
    modem_url = 'http://192.168.100.1/Docsis_system.asp'
    # 设置超时,避免长时间等待
    try:
        page = requests.get(modem_url, timeout=10)
        page.raise_for_status()  # 检查请求是否成功
        soup = BeautifulSoup(page.text, 'html.parser')
    except requests.exceptions.RequestException as e:
        print(f"请求调制解调器页面失败: {e}")
        return

    # 定义要抓取的信道类型和对应的表格标识
    channel_types = [
        {"name": "downstream", "table_label": "Downstream Channel"},
        {"name": "upstream", "table_label": "Upstream Channel"}
    ]

    # 生成带时间戳的CSV文件名,方便日志记录
    timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
    output_file = f"modem_channels_{timestamp}.csv"

    with open(output_file, 'w', newline='', encoding='utf-8') as csvfile:
        writer = None
        first_write = True

        for channel_type in channel_types:
            # 定位对应表格(根据标题文本查找父表格,不同调制解调器页面结构可能略有不同,可根据实际调整)
            table_title = soup.find(text=lambda text: text and channel_type["table_label"] in text)
            if not table_title:
                print(f"未找到{channel_type['name']}信道表格")
                continue

            # 找到标题所在的表格(通常标题是表格的前一个或父元素,这里假设标题在表格内的标题行)
            table = table_title.find_parent("table")
            if not table:
                print(f"无法定位{channel_type['name']}信道表格")
                continue

            # 提取表头
            headers = []
            header_rows = table.find_all("tr")[0]  # 第一行是表头
            for th in header_rows.find_all("th"):
                headers.append(th.get_text(strip=True))
            # 添加信道类型列,区分上行/下行
            headers.append("Channel_Type")

            # 提取数据行
            data_rows = table.find_all("tr")[1:]  # 跳过表头行
            for row in data_rows:
                row_data = []
                for td in row.find_all("td"):
                    row_data.append(td.get_text(strip=True))
                # 添加信道类型标识
                row_data.append(channel_type["name"])

                # 初始化CSV writer,第一次写入时写表头
                if first_write:
                    writer = csv.writer(csvfile)
                    writer.writerow(headers)
                    first_write = False
                writer.writerow(row_data)

    print(f"信道数据已成功导出到文件: {output_file}")

if __name__ == "__main__":
    fetch_modem_channel_data()

关键逻辑说明

  • 请求与异常处理:添加了超时设置和请求异常捕获,避免因调制解调器离线或网络问题导致程序崩溃。
  • 表格定位逻辑:通过表格标题文本查找目标表格,适配大多数Docsis调制解调器的页面结构。如果你的调制解调器页面结构特殊,可以打开http://192.168.100.1/Docsis_system.asp,用浏览器开发者工具查看表格的id/class,调整定位逻辑(比如改用soup.find("table", id="downstream-table"))。
  • CSV标准化处理:
    • 自动生成带时间戳的文件名,方便按时间归档日志。
    • 新增Channel_Type列,明确区分上行/下行数据。
    • 统一处理表头和数据行的提取,确保输出格式符合CSV规范。
  • 可扩展性:如果需要抓取更多指标(比如信号强度、信噪比),只需要调整表头和数据行的提取逻辑即可。

注意事项

  • 确保你的设备和调制解调器在同一局域网内,能正常访问192.168.100.1。
  • 部分调制解调器可能需要登录后才能查看信道数据,如果遇到403/跳转登录页的情况,需要补充模拟登录的逻辑(比如携带Cookie或提交登录表单)。

内容的提问来源于stack exchange,提问作者LB-SDS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:41:29