You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取展会参展商表格指定字段求助

问题整理

我是网络爬虫初学者,耗费数小时尝试解析某即将举办展会的参展商信息页面表格,希望提取若干指定字段但始终未能实现。

现有代码

profile = requests.get('https://annual.asaecenter.org/profile.cfm?profile_name=exhibitor&master_key=EF74CF1F-95BA-EC11-80F4-EC7F36E6C06A&inv_mast_key=93A17E5D-A46F-F21E-77E4-77B38A3B30EE')
soup = bs(profile.content, 'html.parser')

tds = soup.find_all("td")

print(tds)

代码运行返回片段

[<td style="width: 60%;">
                        Alpharetta Convention and Visitors Bureau
                </td>, <td style="width: 40%; text-align: right;">

                                 

                                Booth 2116
                </td>, <td class="tb-text-left" colspan="2">
<div>
</div>
</td>, <td style="width: 40%;">
                Alpharetta Convention and Visitors Bureau                                                           <br/>
</td>, <td style="width: 60%;" valign="top">
</td>, <td colspan="2"><b>Sales Contact</b><br/>
                Beth Brown<br/>
                Vice President of Sales
                </td>, <td colspan="2">
<a class="bttn-form bttn-form-default" href="javascript:Pops('http://www.awesomealpharetta.com','website',750,650)">
<i aria-hidden="true" class="fa fa-globe fa-lg" title="#xlink_label#"></i><br/>Website
                </a>
</td>, <td colspan="2">
                                Description
                        </td>, <td class="bgcolorw" colspan="2" valign="top">
     Alpharetta, GA has 30 hotels w/ 3,940 + guest rooms, 44,000 sq. ft. conference center, 200+ restaurants, 250+ shops, &amp; 40+ attractions for your attendees
        </td>, <td align="left" class="cellGrad" style="vertical-align:text-top; font-weight:bold; width: 175px">
<label for="TBE573973_4058_EC11_80F3_D9AE4409EDD7ID" id="ROW1780B7E3F-A0FD-41FD-BB77-FD8AD8F6356ELabel">Product Categories<label>
</label></label></td>, <td class="bgcolorw" colspan="1" valign="top">
<a href="/profile.cfm?profile_name=match_exhibitor&amp;answer_key=D6573973-4058-EC11-80F3-D9AE4409EDD7&amp;xtemplate">

期望输出

name = Alpharetta Convention and Visitors Bureau
booth = 2116
url = http://www.awesomealpharetta.com
description = Alpharetta, GA has 30 hotels w/ 3,940 + guest rooms, 44,000 sq. ft. conference center, 200+ restaurants, 250+ shops, & 40+ attractions for your attendees

已知字段XPath

name = //*[@id="exhibitor-profile"]/tbody/tr[1]/td[1]
booth = //*[@id="exhibitor-profile"]/tbody/tr[1]/td[2]
website = //*[@id="exhibitor-profile"]/tbody/tr[3]/td/a
description = //*[@id="ROW1466DD5DF-0695-4D68-B221-4941A5171EAB"]/td
实现方案

直接按元素特征定位提取即可,不需要遍历所有td标签,完整可运行代码如下:

import requests
import re
from bs4 import BeautifulSoup as bs

# 请求页面
profile = requests.get('https://annual.asaecenter.org/profile.cfm?profile_name=exhibitor&master_key=EF74CF1F-95BA-EC11-80F4-EC7F36E6C06A&inv_mast_key=93A17E5D-A46F-F21E-77E4-77B38A3B30EE')
soup = bs(profile.content, 'html.parser')

# 提取展商名称
exhibitor_table = soup.find(id="exhibitor-profile")
first_row = exhibitor_table.find("tr")
first_row_tds = first_row.find_all("td")
name = first_row_tds[0].get_text(strip=True)

# 提取展位号
booth_text = first_row_tds[1].get_text(strip=True)
booth = booth_text.replace("Booth ", "").strip()

# 提取官网地址
website_a = exhibitor_table.find("a", string=re.compile("Website"))
js_content = website_a.get("href")
url = re.search(r"Pops\('(http.*?)','website'", js_content).group(1)

# 提取简介
desc_label = soup.find("td", string=re.compile("Description"))
description = desc_label.find_next_sibling("td").get_text(strip=True)

# 打印结果
print(f"name = {name}")
print(f"booth = {booth}")
print(f"url = {url}")
print(f"description = {description}")

代码说明

  • 定位逻辑完全匹配你给出的XPath路径,复用你已经安装的requests和bs4库即可,不需要额外依赖
  • 官网地址从javascript:Pops()的调用参数中通过正则提取,因为a标签本身没有直接写跳转链接
  • 所有文本提取都加了strip=True参数,自动清除内容前后多余的换行、空格和缩进

内容的提问来源于stack exchange,提问作者BQuist

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 15:54:19