如何从指定SEC网页提取第一个无可用ID标识的表格?
提取SEC目标页面第一个表格的方法
你已通过表格ID成功提取页面的第二个表格,针对无ID标识的第一个表格,可通过以下方式提取:
方法1:通过位置索引直接获取
利用BeautifulSoup的find_all('table')方法获取页面所有表格的列表,取索引为0的元素即为第一个表格:
import requests from bs4 import BeautifulSoup import pandas as pd url = 'https://www.sec.gov/cgi-bin/own-disp?action=getissuer&CIK=1318605' # 添加合规请求头,规避SEC反爬限制 headers = {'User-Agent': 'Your Contact Info/1.0 (your.email@example.com)'} response = requests.get(url, headers=headers) soup = BeautifulSoup(response.content, 'html.parser') # 提取第一个表格 first_table = soup.find_all('table')[0] first_report = pd.read_html(str(first_table))[0]
方法2:通过表格特征精准定位
若第一个表格具备独特特征(比如特定class属性、表头内容),可基于特征定位。例如查看页面源码后,若第一个表格的class为table-bordered,则使用:
first_table = soup.find('table', {'class': 'table-bordered'})
注意事项
SEC网站存在反爬机制,请求时必须添加包含联系方式的User-Agent头,符合其访问规范,避免被限制访问。
内容的提问来源于stack exchange,提问作者AaronLbk
相关产品推荐
相关产品推荐

