You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup提取网页br标签间的医院邮编并存储到列表

提取<br>标签分隔的邮编解决方案

核心思路是利用英国邮编的固定格式规则,用正则匹配目标区域的文本即可,无需单独处理<br>标签,实现成本更低、适配性更强。

操作步骤

  1. 引入正则表达式模块re,编写英国邮编的匹配规则
  2. 拿到目标div容器后,先用find('h1')提取医院名称
  3. 调用get_text()方法直接提取容器内所有纯文本(自动剔除所有HTML标签,包括<br>)
  4. 用正则从纯文本中搜索符合规则的邮编即可

修改后可直接运行的完整代码

import requests
from bs4 import BeautifulSoup
import re

# 英国标准邮编正则匹配规则
postcode_pattern = re.compile(r'[A-Z]{1,2}\d[A-Z\d]? \d[A-Z]{2}')

url_list = ['http://www.wales.nhs.uk/ourservices/directory/Hospitals/92',
'http://www.wales.nhs.uk/ourservices/directory/Hospitals/62']

result = {'hospital':[],'address':[]}

for url in url_list:
    resp = requests.get(url)
    soup = BeautifulSoup(resp.content, "lxml")
    target_div = soup.find('div',{'style':'width:500px; float:left; '})
    # 提取医院名称
    hos_name = target_div.find('h1').text.strip()
    result['hospital'].append(hos_name)
    # 提取邮编
    all_text = target_div.get_text()
    postcode_match = postcode_pattern.search(all_text)
    # 兼容匹配不到的场景,避免报错
    if postcode_match:
        result['address'].append(postcode_match.group())
    else:
        result['address'].append('')

print(result)

运行输出结果

{'hospital': ['Bronglais General Hospital', 'Glan Clwyd Hospital'],
 'address': ['SY23 1ER','LL18 5UJ']}

如果需要提取除邮编外的其他地址内容,只需调整正则规则或者拆分文本即可,无需处理零散的换行标签。

内容的提问来源于stack exchange,提问作者Stackcans

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 09:06:00