如何爬取无独特类名的股票筛选网站表格数据?
无唯一类名的表格数据爬取解决方案
你之前使用的选择器没有限定元素所在的表格单元格位置,finviz内幕交易页面的.tab-link类会同时作用于股票代码、持有人姓名、SEC文档链接三类元素,所以会匹配到多余内容。
单独提取Ticker的修正代码
你可以通过限定选择器匹配每行第一个单元格内的.tab-link元素,精准过滤掉其他位置的同类元素:
Ticker = [item.text.strip() for item in soup.select('tr td:first-child .tab-link')]
全字段爬取完整方案
你需要爬取多类数据保存到Excel,更稳定的方案是直接按行按列索引提取所有字段,无需依赖元素类名做复杂匹配,也能避免不同字段单独提取出现数据错位的问题:
from bs4 import BeautifulSoup import requests import pandas as pd headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:94.0) Gecko/20100101 Firefox/94.0'} df_headers = ['Ticker' , 'Owner' , 'Relationship' , 'Date' ,'Transaction' , 'Total Shares' , 'SEC Form'] url= "https://finviz.com/insidertrading.ashx" r = requests.get(url, headers=headers) soup = BeautifulSoup(r.content, 'lxml') # 匹配所有数据行,跳过表头 data_rows = soup.select('table.styled-table-insider tr')[1:] result = [] for row in data_rows: cols = row.find_all('td') # 按列索引提取对应字段 ticker = cols[0].text.strip() owner = cols[1].text.strip() relationship = cols[2].text.strip() date = cols[3].text.strip() transaction = cols[4].text.strip() total_shares = cols[5].text.strip() sec_form = cols[6].text.strip() result.append([ticker, owner, relationship, date, transaction, total_shares, sec_form]) # 导出到Excel df = pd.DataFrame(result, columns=df_headers) df.to_excel('finviz内幕交易数据.xlsx', index=False)
内容的提问来源于stack exchange,提问作者seanofdead
相关产品推荐
相关产品推荐

