You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python实现两列网页爬取并输出指定格式表格

Python两列网页数据爬取修正方案

针对你提供的HTML结构,可使用BeautifulSoup库精准提取姓名和公司信息,输出符合预期的两列表格。

实现代码

from bs4 import BeautifulSoup

# 替换为实际爬取的HTML内容
html = '''
<div class="content">
<div class="container">
<div class="row pt-2">
<div class="col pe-1">
<div class="grid-cell p-2">
<a href="united-states_florida/company/met-west-commercial-lender/tom-mchugh-975">
              Tom Mchugh
            </a>
</div>
</div>
<div class="col ps-1">
<div class="grid-cell p-2">
            Company:
            <span>
<a href="united-states_florida/company/met-west-commercial-lender">
                Met West Commercial Lender
                  </a>
</span>
</div>
</div>
</div>
</div>
</div>
'''

soup = BeautifulSoup(html, 'html.parser')

# 定位所有数据行
rows = soup.find_all('div', class_='row pt-2')

# 输出Markdown格式表格
print("| Column A | Column B |")
print("|----------|----------|")
for row in rows:
    # 提取姓名并清理空格
    name = row.find('div', class_='col pe-1').find('a').get_text(strip=True)
    # 提取公司名并清理空格
    company = row.find('div', class_='col ps-1').find('span').find('a').get_text(strip=True)
    print(f"| {name} | {company} |")

关键说明

  • 以row pt-2为单位遍历,确保每组姓名与公司属于同一数据行
  • 用get_text(strip=True)去除文本两端冗余空格,避免输出格式混乱
  • 直接输出Markdown表格,也可将数据存入列表后导出为CSV等格式

内容的提问来源于stack exchange,提问作者Noah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 03:10:30