You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python抓取嵌套表格数据?解决radiofeeds爬取难题

解决RadioFeeds网站嵌套表格爬取问题

问题背景

爬取RadioFeeds网站整理英国电台列表时,目标页面的Stream URLs位于嵌套表格中。使用BeautifulSoup遍历<tr>标签时,会在父表格与子表格间跳转,导致列数(COL_COUNT)波动,出现空Station Name,无法正确获取包含子表格的行数据。

当前代码

from bs4 import BeautifulSoup
from requests import get

url = "http://www.radiofeeds.co.uk/mp3.asp"
page = get(url=url).text
lead = "Listen live online to"
foot = "Have <b>YOUR</b> internet"
start = page.find(lead)
stop = page.find(foot)

soup = BeautifulSoup(page[start:stop], "html.parser")

data = []
table = soup.find("table")
rows = table.find_all("tr")
for row in rows:
    station_name = row.find("a").text
    print(f"STATION_NAME: {station_name}")

    cols = len(row.find_all("td"))
    print(f"COL_COUNT: {cols}")

    print("=====")

代码输出示例

STATION_NAME: 121 Radio
COL_COUNT: 9
=====
STATION_NAME: 

COL_COUNT: 2
=====
STATION_NAME: 10-fi Radio
COL_COUNT: 9
=====
STATION_NAME: 

COL_COUNT: 2
=====
STATION_NAME: 45 Radio
COL_COUNT: 9
=====
STATION_NAME: 

COL_COUNT: 2
=====

解决方案

核心问题是嵌套表格的<tr>被误抓取,导致无效行混入遍历流程。以下两种方法可解决该问题:

方法一:仅遍历父表格的直接子行

通过recursive=False参数,只获取父表格的直接子<tr>,排除子表格内的行:

from bs4 import BeautifulSoup
from requests import get

url = "http://www.radiofeeds.co.uk/mp3.asp"
page = get(url=url).text
lead = "Listen live online to"
foot = "Have <b>YOUR</b> internet"
start = page.find(lead)
stop = page.find(foot)

soup = BeautifulSoup(page[start:stop], "html.parser")

data = []
table = soup.find("table")
# 只获取父表格的直接子tr,不递归查找子表格内的行
rows = table.find_all("tr", recursive=False)
for row in rows:
    # 先判断是否存在a标签,避免报错
    station_a = row.find("a")
    station_name = station_a.text if station_a else "无名称"
    print(f"STATION_NAME: {station_name}")

    cols = len(row.find_all("td"))
    print(f"COL_COUNT: {cols}")

    print("=====")

方法二:通过特征筛选有效行

如果直接子行的方法不生效,可通过列数(有效行的列数为9)或是否存在电台名称的<a>标签来筛选有效行:

from bs4 import BeautifulSoup
from requests import get

url = "http://www.radiofeeds.co.uk/mp3.asp"
page = get(url=url).text
lead = "Listen live online to"
foot = "Have <b>YOUR</b> internet"
start = page.find(lead)
stop = page.find(foot)

soup = BeautifulSoup(page[start:stop], "html.parser")

data = []
table = soup.find("table")
rows = table.find_all("tr")
for row in rows:
    cols = len(row.find_all("td"))
    # 仅处理列数为9的有效行
    if cols != 9:
        continue
    station_a = row.find("a")
    station_name = station_a.text if station_a else "无名称"
    print(f"STATION_NAME: {station_name}")
    print(f"COL_COUNT: {cols}")
    print("=====")

额外优化建议

可以替换解析器为lxml,它处理嵌套HTML结构更稳定,速度也更快。使用前需先安装:

pip install lxml

然后修改BeautifulSoup初始化代码:

soup = BeautifulSoup(page[start:stop], "lxml")

内容的提问来源于stack exchange,提问作者James Geddes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 02:01:14