使用BeautifulSoup爬取展会参展商表格指定字段求助
问题整理
我是网络爬虫初学者,耗费数小时尝试解析某即将举办展会的参展商信息页面表格,希望提取若干指定字段但始终未能实现。
现有代码
profile = requests.get('https://annual.asaecenter.org/profile.cfm?profile_name=exhibitor&master_key=EF74CF1F-95BA-EC11-80F4-EC7F36E6C06A&inv_mast_key=93A17E5D-A46F-F21E-77E4-77B38A3B30EE') soup = bs(profile.content, 'html.parser') tds = soup.find_all("td") print(tds)
代码运行返回片段
[<td style="width: 60%;"> Alpharetta Convention and Visitors Bureau </td>, <td style="width: 40%; text-align: right;"> Booth 2116 </td>, <td class="tb-text-left" colspan="2"> <div> </div> </td>, <td style="width: 40%;"> Alpharetta Convention and Visitors Bureau <br/> </td>, <td style="width: 60%;" valign="top"> </td>, <td colspan="2"><b>Sales Contact</b><br/> Beth Brown<br/> Vice President of Sales </td>, <td colspan="2"> <a class="bttn-form bttn-form-default" href="javascript:Pops('http://www.awesomealpharetta.com','website',750,650)"> <i aria-hidden="true" class="fa fa-globe fa-lg" title="#xlink_label#"></i><br/>Website </a> </td>, <td colspan="2"> Description </td>, <td class="bgcolorw" colspan="2" valign="top"> Alpharetta, GA has 30 hotels w/ 3,940 + guest rooms, 44,000 sq. ft. conference center, 200+ restaurants, 250+ shops, & 40+ attractions for your attendees </td>, <td align="left" class="cellGrad" style="vertical-align:text-top; font-weight:bold; width: 175px"> <label for="TBE573973_4058_EC11_80F3_D9AE4409EDD7ID" id="ROW1780B7E3F-A0FD-41FD-BB77-FD8AD8F6356ELabel">Product Categories<label> </label></label></td>, <td class="bgcolorw" colspan="1" valign="top"> <a href="/profile.cfm?profile_name=match_exhibitor&answer_key=D6573973-4058-EC11-80F3-D9AE4409EDD7&xtemplate">
期望输出
name = Alpharetta Convention and Visitors Bureau booth = 2116 url = http://www.awesomealpharetta.com description = Alpharetta, GA has 30 hotels w/ 3,940 + guest rooms, 44,000 sq. ft. conference center, 200+ restaurants, 250+ shops, & 40+ attractions for your attendees
已知字段XPath
name = //*[@id="exhibitor-profile"]/tbody/tr[1]/td[1] booth = //*[@id="exhibitor-profile"]/tbody/tr[1]/td[2] website = //*[@id="exhibitor-profile"]/tbody/tr[3]/td/a description = //*[@id="ROW1466DD5DF-0695-4D68-B221-4941A5171EAB"]/td
实现方案
直接按元素特征定位提取即可,不需要遍历所有td标签,完整可运行代码如下:
import requests import re from bs4 import BeautifulSoup as bs # 请求页面 profile = requests.get('https://annual.asaecenter.org/profile.cfm?profile_name=exhibitor&master_key=EF74CF1F-95BA-EC11-80F4-EC7F36E6C06A&inv_mast_key=93A17E5D-A46F-F21E-77E4-77B38A3B30EE') soup = bs(profile.content, 'html.parser') # 提取展商名称 exhibitor_table = soup.find(id="exhibitor-profile") first_row = exhibitor_table.find("tr") first_row_tds = first_row.find_all("td") name = first_row_tds[0].get_text(strip=True) # 提取展位号 booth_text = first_row_tds[1].get_text(strip=True) booth = booth_text.replace("Booth ", "").strip() # 提取官网地址 website_a = exhibitor_table.find("a", string=re.compile("Website")) js_content = website_a.get("href") url = re.search(r"Pops\('(http.*?)','website'", js_content).group(1) # 提取简介 desc_label = soup.find("td", string=re.compile("Description")) description = desc_label.find_next_sibling("td").get_text(strip=True) # 打印结果 print(f"name = {name}") print(f"booth = {booth}") print(f"url = {url}") print(f"description = {description}")
代码说明
- 定位逻辑完全匹配你给出的XPath路径,复用你已经安装的
requests和bs4库即可,不需要额外依赖 - 官网地址从
javascript:Pops()的调用参数中通过正则提取,因为a标签本身没有直接写跳转链接 - 所有文本提取都加了
strip=True参数,自动清除内容前后多余的换行、空格和缩进
内容的提问来源于stack exchange,提问作者BQuist
相关产品推荐
相关产品推荐

