如何用BeautifulSoup从非结构化表格行提取无标签包裹的地址
提取无标签包裹的HTML地址内容
问题场景
目标地址「Rose Avenue 33, 4302843 A City」未被任何HTML标签包裹,无法通过常规find()方法精准定位,现有尝试代码未能成功提取该内容。
目标HTML文档
<!DOCTYPE html> <html lang="en"> <head> <meta charset="UTF-8"> <title>Title</title> </head> <body> <table class="novip"> <tr class="novip"> <td class="novip-portrait-picture" rowspan="5"> <a class="novip" href="refer.html">URL</a> </td> <td class="novip-left"> <a class="novip-firmen-name" href="refer.html" target="_top"> John Doe </a> </td> <td class="novip-right" rowspan="2"> <a class="novip" href="refer.html">URL</a> </td> </tr> <tr class="novip"> <td class="novip-left"> <span class="novip-left-titel"> Prof. </span> <span class="novip-left-fachbezeichnung"> Professor for History </span> <br/> Rose Avenue 33, 4302843 A City <br/> Tel: <a>234 23 43244</a> <a class="novip-left-make_appointment-button-active">Booking</a> </td> </tr> </table> </body> </html>
尝试的代码(未成功)
from bs4 import BeautifulSoup import requests # 原代码遗漏requests导入 r = requests.get(url) r.encoding = 'utf8' html_doc = r.text soup = BeautifulSoup(html_doc, features='html5lib') table = [] tables = soup.find_all("table", {"class": "novip"}) for table in tables: rows = table.findChildren('tr') address = rows[1].find('span', 'novip-left-fachbezeichnung').text
解决方案
地址是直接作为文本节点存在于<td class="novip-left">下,位于第二个<span>和第一个<br/>之后,可通过以下几种方式提取:
方法1:利用next_sibling定位文本节点
通过跳过换行和<br/>标签,直接定位目标文本节点:
from bs4 import BeautifulSoup import requests r = requests.get(url) r.encoding = 'utf8' html_doc = r.text soup = BeautifulSoup(html_doc, features='html5lib') tables = soup.find_all("table", {"class": "novip"}) for table in tables: # 定位目标td target_td = table.find_all('tr')[1].find('td', class_='novip-left') # 找到职称span,跳过后续空白和br标签 fach_span = target_td.find('span', class_='novip-left-fachbezeichnung') address_text = fach_span.next_sibling.next_sibling.next_sibling.strip() print(address_text)
方法2:提取所有文本节点后筛选
提取目标td下的所有非空白文本,根据地址特征(包含逗号和数字)筛选:
from bs4 import BeautifulSoup import requests r = requests.get(url) r.encoding = 'utf8' html_doc = r.text soup = BeautifulSoup(html_doc, features='html5lib') tables = soup.find_all("table", {"class": "novip"}) for table in tables: target_td = table.find_all('tr')[1].find('td', class_='novip-left') # 获取所有去除空白的文本内容 all_texts = [text.strip() for text in target_td.stripped_strings] # 筛选符合地址特征的内容 for text in all_texts: if ',' in text and any(char.isdigit() for char in text): address_text = text print(address_text) break
方法3:定位<br/>标签后提取文本
找到第一个<br/>标签,直接获取其后续的文本节点:
from bs4 import BeautifulSoup import requests r = requests.get(url) r.encoding = 'utf8' html_doc = r.text soup = BeautifulSoup(html_doc, features='html5lib') tables = soup.find_all("table", {"class": "novip"}) for table in tables: target_td = table.find_all('tr')[1].find('td', class_='novip-left') # 找到第一个br标签 first_br = target_td.find('br') # 获取br后的文本并去除空白 address_text = first_br.next_sibling.strip() print(address_text)
内容的提问来源于stack exchange,提问作者CarbonaraSenzaPanna
相关产品推荐
相关产品推荐

