BeautifulSoup网页爬取求助:提取指定td标签的地址信息
解决方法:使用BeautifulSoup提取指定表格行的地址内容
没问题,我来帮你搞定这个爬取需求。下面是直接可行的实现方案,包含代码和关键说明:
核心思路
我们需要先定位到目标表格,然后遍历每一行,检查第一个<td>的文本是否是Legal Adress或Actual Adress,如果匹配就提取同行第二个<td>里的地址内容。
完整代码示例
from bs4 import BeautifulSoup # 你的目标HTML片段(如果是从网页获取,先使用requests库获取页面内容) html_content = """ <div class="col-md-8 col-xs-12"> <table class="table table-bordered"> <!-- <tr> <td width="20%">Seller</td> <td>CYBERNET MƏHDUD MƏSULİYYƏTLİ CƏMİYYƏTİ</td> </tr> --> <tr> <td>Legal Adress</td> <td>AZ1025, Baku city, Khatai District, Suleyman Vazirov, house 1170th quarter, 3-4th block, 1st floor</td> </tr> <tr> <td>Actual Adress</td> <td>AZ1025, Baku city, Khatai District, Suleyman Vazirov, house 1170th quarter, 3-4th block, 1st floor</td> </tr> <tr> <td>Description</td> <td> Cybernet LLC is one of the leading companies providing Information Technology services and solutions in Azerbaijan. Working mainly on Software Development, Modeling System Architecture and Building IT Infrastructure company offers innovative software and advanced hardware application, improvement, technical support and integration services for Government, Bank, Telecommunication, Large and Medium-sized Business organizations on Corporate management, finance and automation of business. Cybernet specialized in automation of Government services and creation of electronic services in Azerbaijan.<br/> </td> </tr> <tr> <td width="30%">TIN</td> <td>9900050571</td> </tr> <tr> <td width="20%">Activity Group</td> <td>58290 - Publication of other software products</td> </tr> <tr> <td>E-mail</td> <td>office@cybernet.az</td> </tr> <tr> <td>Phone</td> <td>+994124038963</td> </tr> <tr> <td>Mobile</td> <td>+994505891920</td> </tr> <tr> <td>Facebook</td> <td><a href="https://www.facebook.com/cybernet.az/" target="_blank">https://www.facebook.com/cybernet.az/</a></td> </tr> <tr> <td>Website</td> <td><a href="http://cybernet.az/" target="_blank">http://cybernet.az/</a></td> </tr> </table> </div> """ # 初始化BeautifulSoup解析器 soup = BeautifulSoup(html_content, 'html.parser') # 定义需要匹配的标签文本 target_labels = ["Legal Adress", "Actual Adress"] # 定位到目标表格 target_table = soup.find('table', class_='table table-bordered') # 遍历表格的每一行 for row in target_table.find_all('tr'): # 获取当前行的所有单元格 cells = row.find_all('td') # 确保行内至少有两个单元格(避免无效行) if len(cells) >= 2: # 提取第一个单元格的文本并去除首尾空白 current_label = cells[0].get_text(strip=True) # 检查是否是目标标签 if current_label in target_labels: # 提取对应的地址内容 address_content = cells[1].get_text(strip=True) # 输出结果(也可以存入字典或列表方便后续处理) print(f"{current_label}: {address_content}")
关键细节说明
- 定位表格:使用
find('table', class_='table table-bordered')精准定位到目标表格,避免页面中其他表格干扰。 - 处理文本空白:使用
get_text(strip=True)去除文本周围的空格、换行符,确保标签匹配的准确性。 - 容错处理:判断
len(cells) >= 2,避免遇到只有一个单元格的无效行导致报错。 - 扩展灵活性:把目标标签放在列表
target_labels中,后续需要添加其他标签时直接修改列表即可。
如果是从真实网页爬取,只需把html_content替换为requests.get(url).text(记得先安装并导入requests库)。
内容的提问来源于stack exchange,提问作者iftixar kerimov
相关产品推荐
相关产品推荐

