如何用Python批量抓取网页中律师个人链接的详情数据?
批量抓取Chambers亚太区律师信息的实现方案
核心思路
先从列表页批量提取所有律师的基础信息(姓名、所属律所、排名)和个人主页链接,再通过循环遍历这些链接批量抓取详情页内容,同时通过优化请求方式提升效率。
步骤1:批量提取列表页的律师链接与基础信息
一次性爬取列表页的所有律师条目,把需要的基础信息和详情页链接收集起来,避免重复请求列表页。以下是Python示例代码(需根据页面实际HTML结构调整选择器):
import requests from bs4 import BeautifulSoup def extract_lawyer_list(): list_url = "https://chambers.com/all-lawyers-asia-pacific-8" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } resp = requests.get(list_url, headers=headers) soup = BeautifulSoup(resp.text, "html.parser") lawyer_data = [] # 替换为页面实际的律师条目选择器 for item in soup.select(".search-result-item"): name = item.select_one(".lawyer-name").get_text(strip=True) firm = item.select_one(".lawyer-firm").get_text(strip=True) ranking = item.select_one(".rank-badge").get_text(strip=True) profile_href = item.select_one(".profile-link")["href"] # 补全绝对链接 profile_url = f"https://chambers.com{profile_href}" if not profile_href.startswith("http") else profile_href lawyer_data.append({ "name": name, "firm": firm, "ranking": ranking, "profile_url": profile_url }) return lawyer_data
步骤2:循环批量抓取详情页
拿到所有律师的详情页链接后,用循环遍历请求每个链接,提取详情页信息。加入异常处理避免单个请求失败中断整个流程,同时加入延迟避免触发反爬:
import csv import time from concurrent.futures import ThreadPoolExecutor def scrape_profile(lawyer): headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } try: resp = requests.get(lawyer["profile_url"], headers=headers) resp.raise_for_status() soup = BeautifulSoup(resp.text, "html.parser") # 提取详情页信息,示例为联系方式和执业领域,需根据页面调整 contact = soup.select_one(".lawyer-contact-info").get_text(strip=True) if soup.select_one(".lawyer-contact-info") else "无" practice_areas = [area.get_text(strip=True) for area in soup.select(".practice-area-tag")] lawyer.update({ "contact": contact, "practice_areas": ", ".join(practice_areas) }) time.sleep(0.8) # 延迟控制请求频率 return lawyer except Exception as e: print(f"抓取失败:{lawyer['name']} - {str(e)}") # 记录失败链接以便重试 with open("failed_profiles.txt", "a", encoding="utf-8") as f: f.write(f"{lawyer['profile_url']}\n") return None def batch_scrape(lawyer_list): # 用多线程提升效率,线程数根据网站反爬强度调整 with ThreadPoolExecutor(max_workers=15) as executor: results = executor.map(scrape_profile, lawyer_list) # 过滤失败的条目,写入CSV valid_results = [res for res in results if res] with open("chambers_lawyers.csv", "w", newline="", encoding="utf-8") as f: writer = csv.DictWriter(f, fieldnames=["name", "firm", "ranking", "profile_url", "contact", "practice_areas"]) writer.writeheader() writer.writerows(valid_results) # 执行流程 if __name__ == "__main__": lawyers = extract_lawyer_list() batch_scrape(lawyers)
效率优化建议
- 多线程/多进程:用
ThreadPoolExecutor替代单线程循环,同时处理多个请求,能把抓取速度提升数倍(注意线程数不要超过20,避免被封)。 - 请求缓存:用
requests-cache库缓存已请求的页面,中途中断后无需重复请求已完成的页面。 - 代理轮换:如果遇到IP限制,使用代理池轮换IP,避免被网站封禁。
注意事项
- 先查看网站的
robots.txt文件,确认允许抓取这类内容,避免违反网站规则。 - 模拟正常用户行为:使用真实的User-Agent,避免固定间隔请求,可随机调整延迟时间(比如0.5-1.5秒)。
- 数据分批存储:如果数据量太大,分批写入文件或数据库,避免内存占用过高。
内容的提问来源于stack exchange,提问作者Psychedelique23
相关产品推荐
相关产品推荐

