如何用BeautifulSoup解析含Microsoft Excel Schema的HTML表格文档?
马来西亚中学爬虫表格解析问题排查与修复
你的问题主要出在三个核心点:
1. 子页面HTML结构非标准
目标.htm页面是Excel导出的HTML文件,没有标准的<tbody>标签,表格行<tr>直接嵌套在<table>标签下,所以soup2.find('tbody')必然返回None。同时页面带有Excel专属命名空间,会干扰常规HTML解析逻辑。
2. 请求头缺失导致内容异常
requests默认请求头会被网站识别为爬虫,可能返回不完整或异常内容,直接导致解析失败。需要添加模拟浏览器的User-Agent请求头。
3. 解析逻辑未匹配目标数据位置
子页面的学校信息分布在表格的特定列中,原代码未定位到正确的<tr>和<td>节点,无法提取名称、邮箱、官网等字段。
修复后的完整代码
import requests from bs4 import BeautifulSoup import csv def write_to_csv(file_name, data): # 用csv模块规范写入,避免手动拼接的格式错误 with open(file_name, "a", newline="", encoding="utf-8") as f: writer = csv.writer(f) writer.writerow(data) def get_soup(url, parser="lxml"): # 统一请求逻辑,添加浏览器请求头 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(url, headers=headers) response.encoding = "utf-8" # 确保中文编码正确 return BeautifulSoup(response.text, parser) def data_fetch(url): soup = get_soup(url) links = soup.find(class_='entry-content').find_all('a') # 初始化CSV表头 write_to_csv("malaysia_schools.csv", ["学校名称", "邮箱", "官网"]) for link in links: web = link.get('href') if not web.endswith(".htm"): continue # 过滤非目标子页面链接 soup2 = get_soup(web) # 直接定位表格(Excel导出的HTML无tbody,tr直接属于table) table = soup2.find("table") if not table: print(f"无法找到表格: {web}") continue rows = table.find_all("tr") # 跳过表头行(根据页面实际结构调整) for row in rows[1:]: cols = row.find_all("td") if len(cols) < 5: continue # 跳过内容不全的行 # 按列索引提取数据(示例页面中名称在第1列,邮箱第4列,官网第5列) school_name = cols[0].get_text(strip=True) email = cols[3].get_text(strip=True) website = cols[4].get_text(strip=True) write_to_csv("malaysia_schools.csv", [school_name, email, website]) print(f"已完成页面爬取: {web}") # break # 测试时保留该行,正式爬取删除 def main(): url = "https://myschoolchildren.com/list-of-all-secondary-schools-in-malaysia/#.YzWrtXZBy3A" data_fetch(url) if __name__ == "__main__": main()
关键修改说明
- 统一请求逻辑,添加
User-Agent模拟浏览器,避免被网站拦截。 - 使用
csv模块写入数据,避免手动拼接字符串导致的CSV格式错误。 - 适配Excel导出HTML的结构,直接查找
<table>标签而非不存在的<tbody>。 - 根据子页面的列位置精准提取目标字段,确保数据准确性。
内容的提问来源于stack exchange,提问作者theycallmepix
相关产品推荐
相关产品推荐

