You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup提取网页p标签末尾的地址类文本

解决方案

实现思路

你需要提取的美国地址有明确的格式特征:末尾固定为「2位大写州缩写 + 空格 + 5位数字邮编」,且都位于对应p标签文本的最后一行,因此可以先拆分p标签的文本行取末尾行,再通过正则校验筛选地址即可。

完整可运行代码

import requests
from bs4 import BeautifulSoup
import re

# 匹配美国地址末尾特征的正则规则
address_pattern = re.compile(r'[A-Z]{2}\s\d{5}$')

url = 'https://www.housebeautiful.com/lifestyle/g26859396/movie-homes-you-can-visit/'
soup = BeautifulSoup(requests.get(url).content, 'lxml')

collected_addresses = []
for p_tag in soup.select('p'):
    text = p_tag.text.strip()
    if not text:
        continue
    # 按换行拆分文本,取最后一行做校验
    last_line = text.split('\n')[-1].strip()
    if address_pattern.search(last_line):
        collected_addresses.append(last_line)
        print(last_line)

补充说明

  • 上述代码会把所有符合格式的地址存入collected_addresses列表,同时直接打印输出结果
  • 如果遇到地址不在p标签最后一行的场景,可以直接在整段p文本中搜索符合地址格式的内容,修改匹配逻辑即可:
for p_tag in soup.select('p'):
    text = p_tag.text.strip()
    addr_match = re.search(r'[\w\s,]+[A-Z]{2} \d{5}', text)
    if addr_match:
        print(addr_match.group().strip())

内容的提问来源于stack exchange,提问作者Stackcans

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 07:45:02