如何用Python爬取Adviserinfo.sec.gov的理财顾问数据?
问题诊断与修复步骤
1. 请求方式错误:改用POST提交搜索
目标网站的搜索功能实际通过POST请求提交表单数据,而非代码中使用的GET请求。直接用GET请求会返回无效页面,导致无法获取搜索结果。
2. 添加必要请求头,绕过基础反爬
网站会检查请求的User-Agent等标识,缺少时会拒绝服务或返回空内容。需要在请求中添加模拟浏览器的请求头。
3. 修正页面元素选择器
原代码中的HTML元素类名已失效,需根据当前页面结构调整选择器。
修复后的完整代码
import requests from bs4 import BeautifulSoup def search_adviser_by_name(first_name, last_name): # 实际搜索提交的POST端点 search_url = "https://adviserinfo.sec.gov/IAPD/Individual/Search/Submit" # 模拟浏览器请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Content-Type': 'application/x-www-form-urlencoded', 'Referer': 'https://adviserinfo.sec.gov/IAPD/Individual/Search/Search' } # POST表单数据 search_data = { 'ADVANCED': 'true', 'FIND_BY_NAME': 'true', 'INDIVIDUAL_NAME': f"{first_name} {last_name}", 'resultsPerPage': '10', 'submit': 'Search' } response = requests.post(search_url, data=search_data, headers=headers) if response.status_code != 200: print(f"搜索请求失败,状态码:{response.status_code}") return None soup = BeautifulSoup(response.content, 'html.parser') # 修正搜索结果选择器 search_results = soup.find_all('div', class_='row individual-result') for result in search_results: name_elem = result.find('h3', class_='individual-name') if not name_elem: continue full_name = name_elem.text.strip() if first_name.lower() in full_name.lower() and last_name.lower() in full_name.lower(): # 获取详情页链接 detail_link = result.find('a', class_='btn btn-primary')['href'] adviser_url = "https://adviserinfo.sec.gov" + detail_link return get_adviser_info(adviser_url, headers) print("未找到匹配的顾问信息") return None def get_adviser_info(adviser_url, headers): response = requests.get(adviser_url, headers=headers) if response.status_code != 200: print(f"详情页请求失败,状态码:{response.status_code}") return None soup = BeautifulSoup(response.content, 'html.parser') # 修正详情页元素选择器 name = soup.find('h1', class_='individual-header-name').text.strip() firm = soup.find('a', class_='firm-name').text.strip() crd_number = soup.find('span', class_='crd-number').text.strip().replace('CRD#:', '').strip() return { 'Name': name, 'Firm': firm, 'CRD Number': crd_number } # 示例使用 first_name = 'Kelly' last_name = 'Demers' info = search_adviser_by_name(first_name, last_name) if info: print(info) else: print("Failed to retrieve adviser information.")
额外注意事项
- 若仍无法获取数据,需检查网站是否有JavaScript动态加载内容,此时需改用
selenium等工具模拟浏览器行为。 - 请勿高频请求,避免触发网站反爬机制导致IP被封禁。
内容的提问来源于stack exchange,提问作者NewToPython
相关产品推荐
相关产品推荐

