如何在Python中遍历非表格元素并抓取数据到DataFrame?
问题描述
需要抓取网站https://www.nhlpa.com/the-pa/certified-agents?range=A-Z上的NHLPA认证代理信息,包括姓名、头像URL、公司、地址、教育信息,最终整理为表格。现有代码无法正确获取内容组件内的数据,代码如下:
r=requests.get(url) soup=BeautifulSoup(r.text, 'html5lib') table = soup.find_all('div', attrs = {'class':'col-lg-6 agent'}) for a in table: if a.find('div', attrs = {'headshot'}): headshot_url=a.find('div', attrs = {'headshot'}).img
解决方案
核心问题修正
- 添加请求头:目标网站会拦截无标识的请求,需添加
User-Agent模拟浏览器访问。 - 修正元素定位语法:原代码中
attrs={'headshot'}写法错误,应指定class属性为headshot;同时需逐个定位所需字段的容器。
完整代码示例
import requests from bs4 import BeautifulSoup import pandas as pd url = "https://www.nhlpa.com/the-pa/certified-agents?range=A-Z" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } # 发起请求并解析页面 response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'html5lib') agent_containers = soup.find_all('div', class_='col-lg-6 agent') # 存储数据的列表 agent_data = [] for container in agent_containers: # 提取头像URL(处理相对路径) headshot_div = container.find('div', class_='headshot') headshot_url = headshot_div.img['src'] if headshot_div and headshot_div.img else None if headshot_url and not headshot_url.startswith('http'): headshot_url = f"https://www.nhlpa.com{headshot_url}" # 提取姓名 name = container.find('div', class_='agent-name').get_text(strip=True) if container.find('div', class_='agent-name') else None # 提取公司 company = container.find('div', class_='agent-company').get_text(strip=True) if container.find('div', class_='agent-company') else None # 提取地址(处理换行和空格) address = container.find('div', class_='agent-address').get_text(strip=True).replace('\n', ', ') if container.find('div', class_='agent-address') else None # 提取教育信息 education = container.find('div', class_='agent-education').get_text(strip=True) if container.find('div', class_='agent-education') else None # 将数据添加到列表 agent_data.append({ '姓名': name, '头像URL': headshot_url, '公司': company, '地址': address, '教育信息': education }) # 转换为DataFrame并输出为表格 df = pd.DataFrame(agent_data) print(df.to_markdown(index=False))
说明
- 代码中通过
class_属性直接定位元素,比attrs写法更简洁; - 对所有字段添加了存在性判断,避免因部分代理信息缺失导致报错;
- 头像URL处理了相对路径问题,拼接为完整的绝对URL;
- 使用
pandas将数据转换为Markdown格式表格,便于直接查看或导出。
内容的提问来源于stack exchange,提问作者user3592894
相关产品推荐
相关产品推荐

