You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python的Beautiful Soup提取HTML表格中的申请截止日期文本

用BeautifulSoup提取HTML表格奖学金截止日期的实现方案

原有代码问题说明

  • 变量名不匹配:你将查询到的目标表格存储在External_Awards变量中,后续循环却调用了未定义的table变量,运行会直接报错
  • 选择器逻辑错误:<br>是换行标签,本身不包含文本内容,你需要的截止日期属于<br>标签后紧跟的文本节点,无法直接通过.text读取<br>的内容获取

修复后可运行代码

import requests
import bs4 as bs
import re

# 请求目标页面
url = "http://www.cityu.edu.hk/sds/web/studentlife_scholarships_awards.shtml"
# 加请求头避免被反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}
resp = requests.get(url, headers=headers)
resp.encoding = 'utf-8'

# 解析页面
soup = bs.BeautifulSoup(resp.text, 'html.parser')
# 定位所有目标表格
target_tables = soup.find_all('table', class_='marginBottom10')

application_deadline = []
# 编译正则匹配指定格式的日期
date_pattern = re.compile(r'\d{1,2}\s*[ap]\.m\.,\s*\d{1,2}\s+\w+\s+\d{4}\s+\([^)]+\)', re.IGNORECASE)

# 遍历所有表格、行、单元格提取日期
for table in target_tables:
    for row in table.find_all('tr'):
        cells = row.find_all('td')
        for cell in cells:
            cell_text = cell.get_text(strip=False)
            # 查找所有符合格式的日期
            matched_dates = date_pattern.findall(cell_text)
            for d in matched_dates:
                # 去除多余空白符
                clean_date = ' '.join(d.split())
                if clean_date not in application_deadline:
                    application_deadline.append(clean_date)

# 输出结果
print("提取到的所有申请截止日期:")
for idx, date in enumerate(application_deadline, 1):
    print(f"{idx}. {date}")

代码逻辑说明

  • 新增User-Agent请求头,避免被网站反爬策略拦截
  • 用正则匹配固定格式的截止日期,不需要依赖页面的换行标签结构,容错率更高
  • 遍历所有目标表格的单元格,清洗提取到的日期文本后自动去重存储

内容的提问来源于stack exchange,提问作者Jamo T

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 15:36:03