如何用Python抓取嵌套表格数据?解决radiofeeds爬取难题
解决RadioFeeds网站嵌套表格爬取问题
问题背景
爬取RadioFeeds网站整理英国电台列表时,目标页面的Stream URLs位于嵌套表格中。使用BeautifulSoup遍历<tr>标签时,会在父表格与子表格间跳转,导致列数(COL_COUNT)波动,出现空Station Name,无法正确获取包含子表格的行数据。
当前代码
from bs4 import BeautifulSoup from requests import get url = "http://www.radiofeeds.co.uk/mp3.asp" page = get(url=url).text lead = "Listen live online to" foot = "Have <b>YOUR</b> internet" start = page.find(lead) stop = page.find(foot) soup = BeautifulSoup(page[start:stop], "html.parser") data = [] table = soup.find("table") rows = table.find_all("tr") for row in rows: station_name = row.find("a").text print(f"STATION_NAME: {station_name}") cols = len(row.find_all("td")) print(f"COL_COUNT: {cols}") print("=====")
代码输出示例
STATION_NAME: 121 Radio COL_COUNT: 9 ===== STATION_NAME: COL_COUNT: 2 ===== STATION_NAME: 10-fi Radio COL_COUNT: 9 ===== STATION_NAME: COL_COUNT: 2 ===== STATION_NAME: 45 Radio COL_COUNT: 9 ===== STATION_NAME: COL_COUNT: 2 =====
解决方案
核心问题是嵌套表格的<tr>被误抓取,导致无效行混入遍历流程。以下两种方法可解决该问题:
方法一:仅遍历父表格的直接子行
通过recursive=False参数,只获取父表格的直接子<tr>,排除子表格内的行:
from bs4 import BeautifulSoup from requests import get url = "http://www.radiofeeds.co.uk/mp3.asp" page = get(url=url).text lead = "Listen live online to" foot = "Have <b>YOUR</b> internet" start = page.find(lead) stop = page.find(foot) soup = BeautifulSoup(page[start:stop], "html.parser") data = [] table = soup.find("table") # 只获取父表格的直接子tr,不递归查找子表格内的行 rows = table.find_all("tr", recursive=False) for row in rows: # 先判断是否存在a标签,避免报错 station_a = row.find("a") station_name = station_a.text if station_a else "无名称" print(f"STATION_NAME: {station_name}") cols = len(row.find_all("td")) print(f"COL_COUNT: {cols}") print("=====")
方法二:通过特征筛选有效行
如果直接子行的方法不生效,可通过列数(有效行的列数为9)或是否存在电台名称的<a>标签来筛选有效行:
from bs4 import BeautifulSoup from requests import get url = "http://www.radiofeeds.co.uk/mp3.asp" page = get(url=url).text lead = "Listen live online to" foot = "Have <b>YOUR</b> internet" start = page.find(lead) stop = page.find(foot) soup = BeautifulSoup(page[start:stop], "html.parser") data = [] table = soup.find("table") rows = table.find_all("tr") for row in rows: cols = len(row.find_all("td")) # 仅处理列数为9的有效行 if cols != 9: continue station_a = row.find("a") station_name = station_a.text if station_a else "无名称" print(f"STATION_NAME: {station_name}") print(f"COL_COUNT: {cols}") print("=====")
额外优化建议
可以替换解析器为lxml,它处理嵌套HTML结构更稳定,速度也更快。使用前需先安装:
pip install lxml
然后修改BeautifulSoup初始化代码:
soup = BeautifulSoup(page[start:stop], "lxml")
内容的提问来源于stack exchange,提问作者James Geddes
相关产品推荐
相关产品推荐

