You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中遍历非表格元素并抓取数据到DataFrame?

问题描述

需要抓取网站https://www.nhlpa.com/the-pa/certified-agents?range=A-Z上的NHLPA认证代理信息,包括姓名、头像URL、公司、地址、教育信息,最终整理为表格。现有代码无法正确获取内容组件内的数据,代码如下:

r=requests.get(url)
soup=BeautifulSoup(r.text, 'html5lib')
table = soup.find_all('div', attrs = {'class':'col-lg-6 agent'}) 
for a in table:
    if a.find('div', attrs = {'headshot'}):
        headshot_url=a.find('div', attrs = {'headshot'}).img
解决方案

核心问题修正

  1. 添加请求头:目标网站会拦截无标识的请求,需添加User-Agent模拟浏览器访问。
  2. 修正元素定位语法:原代码中attrs={'headshot'}写法错误,应指定class属性为headshot;同时需逐个定位所需字段的容器。

完整代码示例

import requests
from bs4 import BeautifulSoup
import pandas as pd

url = "https://www.nhlpa.com/the-pa/certified-agents?range=A-Z"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

# 发起请求并解析页面
response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, 'html5lib')
agent_containers = soup.find_all('div', class_='col-lg-6 agent')

# 存储数据的列表
agent_data = []

for container in agent_containers:
    # 提取头像URL(处理相对路径)
    headshot_div = container.find('div', class_='headshot')
    headshot_url = headshot_div.img['src'] if headshot_div and headshot_div.img else None
    if headshot_url and not headshot_url.startswith('http'):
        headshot_url = f"https://www.nhlpa.com{headshot_url}"
    
    # 提取姓名
    name = container.find('div', class_='agent-name').get_text(strip=True) if container.find('div', class_='agent-name') else None
    
    # 提取公司
    company = container.find('div', class_='agent-company').get_text(strip=True) if container.find('div', class_='agent-company') else None
    
    # 提取地址(处理换行和空格)
    address = container.find('div', class_='agent-address').get_text(strip=True).replace('\n', ', ') if container.find('div', class_='agent-address') else None
    
    # 提取教育信息
    education = container.find('div', class_='agent-education').get_text(strip=True) if container.find('div', class_='agent-education') else None
    
    # 将数据添加到列表
    agent_data.append({
        '姓名': name,
        '头像URL': headshot_url,
        '公司': company,
        '地址': address,
        '教育信息': education
    })

# 转换为DataFrame并输出为表格
df = pd.DataFrame(agent_data)
print(df.to_markdown(index=False))

说明

  • 代码中通过class_属性直接定位元素,比attrs写法更简洁;
  • 对所有字段添加了存在性判断,避免因部分代理信息缺失导致报错;
  • 头像URL处理了相对路径问题,拼接为完整的绝对URL;
  • 使用pandas将数据转换为Markdown格式表格,便于直接查看或导出。

内容的提问来源于stack exchange,提问作者user3592894

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 10:10:15