You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup从非结构化表格行提取无标签包裹的地址

提取无标签包裹的HTML地址内容

问题场景

目标地址「Rose Avenue 33, 4302843 A City」未被任何HTML标签包裹,无法通过常规find()方法精准定位,现有尝试代码未能成功提取该内容。

目标HTML文档

<!DOCTYPE html>
<html lang="en">
<head>
    <meta charset="UTF-8">
    <title>Title</title>
</head>
<body>
<table class="novip">
    <tr class="novip">
        <td class="novip-portrait-picture"
            rowspan="5">
            <a class="novip" href="refer.html">URL</a>
        </td>
        <td class="novip-left">
            <a class="novip-firmen-name"
               href="refer.html"
               target="_top">
                John Doe
            </a>
        </td>
        <td class="novip-right"
            rowspan="2">
            <a class="novip" href="refer.html">URL</a>
        </td>
    </tr>
    <tr class="novip">
        <td class="novip-left">
            <span class="novip-left-titel">
              Prof.
            </span>
            <span class="novip-left-fachbezeichnung">
              Professor for History
            </span>
            <br/>
            Rose Avenue 33, 4302843 A City
            <br/>
            Tel:&nbsp;<a>234 23 43244</a>
            &nbsp;&nbsp;
            <a class="novip-left-make_appointment-button-active">Booking</a>
            &nbsp;&nbsp;
        </td>
    </tr>

</table>

</body>
</html>

尝试的代码(未成功)

from bs4 import BeautifulSoup
import requests  # 原代码遗漏requests导入

r = requests.get(url)
r.encoding = 'utf8'
html_doc = r.text
soup = BeautifulSoup(html_doc, features='html5lib')
table = []

tables = soup.find_all("table", {"class": "novip"})

for table in tables:
    rows = table.findChildren('tr')
    
    address = rows[1].find('span', 'novip-left-fachbezeichnung').text

解决方案

地址是直接作为文本节点存在于<td class="novip-left">下,位于第二个<span>和第一个<br/>之后,可通过以下几种方式提取:

方法1:利用next_sibling定位文本节点

通过跳过换行和<br/>标签,直接定位目标文本节点:

from bs4 import BeautifulSoup
import requests

r = requests.get(url)
r.encoding = 'utf8'
html_doc = r.text
soup = BeautifulSoup(html_doc, features='html5lib')

tables = soup.find_all("table", {"class": "novip"})
for table in tables:
    # 定位目标td
    target_td = table.find_all('tr')[1].find('td', class_='novip-left')
    # 找到职称span,跳过后续空白和br标签
    fach_span = target_td.find('span', class_='novip-left-fachbezeichnung')
    address_text = fach_span.next_sibling.next_sibling.next_sibling.strip()
    print(address_text)

方法2:提取所有文本节点后筛选

提取目标td下的所有非空白文本,根据地址特征(包含逗号和数字)筛选:

from bs4 import BeautifulSoup
import requests

r = requests.get(url)
r.encoding = 'utf8'
html_doc = r.text
soup = BeautifulSoup(html_doc, features='html5lib')

tables = soup.find_all("table", {"class": "novip"})
for table in tables:
    target_td = table.find_all('tr')[1].find('td', class_='novip-left')
    # 获取所有去除空白的文本内容
    all_texts = [text.strip() for text in target_td.stripped_strings]
    # 筛选符合地址特征的内容
    for text in all_texts:
        if ',' in text and any(char.isdigit() for char in text):
            address_text = text
            print(address_text)
            break

方法3:定位<br/>标签后提取文本

找到第一个<br/>标签,直接获取其后续的文本节点:

from bs4 import BeautifulSoup
import requests

r = requests.get(url)
r.encoding = 'utf8'
html_doc = r.text
soup = BeautifulSoup(html_doc, features='html5lib')

tables = soup.find_all("table", {"class": "novip"})
for table in tables:
    target_td = table.find_all('tr')[1].find('td', class_='novip-left')
    # 找到第一个br标签
    first_br = target_td.find('br')
    # 获取br后的文本并去除空白
    address_text = first_br.next_sibling.strip()
    print(address_text)

内容的提问来源于stack exchange,提问作者CarbonaraSenzaPanna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 23:33:22