如何用Python BeautifulSoup抓取无ID/Class表格并转为DataFrame?
定位目标表格并抓取数据的解决方案
目标表格没有专属ID或类名,我们可以通过它上方的标题文本"Possible Bogus Routes"来定位,具体操作如下:
- 先找到包含该标题的
<h3>元素 - 从这个标题元素出发,定位它后面紧邻的表格
- 遍历表格数据行,提取每列内容整理成字典列表,最后转为DataFrame
完整代码如下:
import requests from bs4 import BeautifulSoup import pandas as pd URL = "https://www.cidr-report.org/as2.0/" page = requests.get(URL) soup = BeautifulSoup(page.content, "html.parser") # 定位"Possible Bogus Routes"标题 bogus_title = soup.find("h3", string="Possible Bogus Routes") # 获取标题后方的目标表格 bogus_table = bogus_title.find_next("table") data_list = [] # 遍历表格行(跳过表头行) for row in bogus_table.find_all("tr")[1:]: cells = row.find_all("td") # 提取各列数据并整理为字典 data_item = { "prefix": cells[0].text.strip(), "origin": cells[1].text.strip(), "description": cells[2].text.strip(), "unallocated": cells[3].text.strip() } data_list.append(data_item) # 转为DataFrame df = pd.DataFrame(data_list) # 打印字典列表格式的结果 print(data_list) # 打印DataFrame预览 print(df.head())
关键说明:
soup.find("h3", string="Possible Bogus Routes"):精准匹配标题文本,锁定目标表格的前置标记bogus_title.find_next("table"):从标题元素直接定位后续的目标表格,避免抓取无关表格- 使用
strip()去除文本前后空白,保证数据整洁
内容的提问来源于stack exchange,提问作者FruitSnacks19
相关产品推荐
相关产品推荐

