You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python的BeautifulSoup提取指定URL网页Location板块下的地址文本

解决方法

你之前的方案依赖文本拆分后的索引和逗号分隔规则,只要字段内容含逗号或者页面字段顺序变化就会失效,直接通过HTML节点结构定位可以完全规避这个问题,步骤如下:

  1. 先定位到Location版块对应的卡片容器
  2. 在容器内匹配Address字段对应的标签,直接提取对应值

以下是可直接运行的代码:

from bs4 import BeautifulSoup
from urllib.request import urlopen, Request

# 加请求头避免站点拦截
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'}

# 两个测试链接都可以正常适配
url = "https://rtk.rjifuture.org/rmp/facility/100000083214"
req = Request(url=url, headers=headers)
html = urlopen(req)
soup = BeautifulSoup(html, "html.parser")

# 定位Location板块
location_card = None
for card in soup.find_all('div', class_='card'):
    header = card.find('div', class_='card-header')
    if header and 'Location' in header.get_text(strip=True):
        location_card = card
        break

# 提取Address字段对应值
address = ""
if location_card:
    dt_list = location_card.find_all('dt')
    for dt in dt_list:
        if dt.get_text(strip=True) == 'Address':
            address = dt.find_next_sibling('dd').get_text(strip=True)
            break

print(address)

上述代码不受字段顺序、地址内含逗号的影响,你提供的两个测试链接都可以正常提取到地址。

内容的提问来源于stack exchange,提问作者color_blue

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 23:21:03