如何使用Python的Beautiful Soup提取HTML表格中的申请截止日期文本
用BeautifulSoup提取HTML表格奖学金截止日期的实现方案
原有代码问题说明
- 变量名不匹配:你将查询到的目标表格存储在
External_Awards变量中,后续循环却调用了未定义的table变量,运行会直接报错 - 选择器逻辑错误:
<br>是换行标签,本身不包含文本内容,你需要的截止日期属于<br>标签后紧跟的文本节点,无法直接通过.text读取<br>的内容获取
修复后可运行代码
import requests import bs4 as bs import re # 请求目标页面 url = "http://www.cityu.edu.hk/sds/web/studentlife_scholarships_awards.shtml" # 加请求头避免被反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } resp = requests.get(url, headers=headers) resp.encoding = 'utf-8' # 解析页面 soup = bs.BeautifulSoup(resp.text, 'html.parser') # 定位所有目标表格 target_tables = soup.find_all('table', class_='marginBottom10') application_deadline = [] # 编译正则匹配指定格式的日期 date_pattern = re.compile(r'\d{1,2}\s*[ap]\.m\.,\s*\d{1,2}\s+\w+\s+\d{4}\s+\([^)]+\)', re.IGNORECASE) # 遍历所有表格、行、单元格提取日期 for table in target_tables: for row in table.find_all('tr'): cells = row.find_all('td') for cell in cells: cell_text = cell.get_text(strip=False) # 查找所有符合格式的日期 matched_dates = date_pattern.findall(cell_text) for d in matched_dates: # 去除多余空白符 clean_date = ' '.join(d.split()) if clean_date not in application_deadline: application_deadline.append(clean_date) # 输出结果 print("提取到的所有申请截止日期:") for idx, date in enumerate(application_deadline, 1): print(f"{idx}. {date}")
代码逻辑说明
- 新增
User-Agent请求头,避免被网站反爬策略拦截 - 用正则匹配固定格式的截止日期,不需要依赖页面的换行标签结构,容错率更高
- 遍历所有目标表格的单元格,清洗提取到的日期文本后自动去重存储
内容的提问来源于stack exchange,提问作者Jamo T
相关产品推荐
相关产品推荐

