You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中遍历HTML表格行并实现网页数据抓取?

Python遍历HTML表格行并提取学生总数的解决方案

原代码的问题

你当前代码的核心错误是元素定位逻辑错误:soup.find_all('a', {'class': 'table table-striped'})试图寻找带有table table-striped类的<a>标签,但实际上这个类属于页面中的<table>元素,并非链接标签,所以根本找不到目标详情链接。

修正后的实现步骤

  1. 定位页面中正确的表格元素
  2. 遍历表格的数据行(跳过表头行)
  3. 提取每行中的详情页链接
  4. 访问详情页,定位并提取学生总数
  5. 整理数据并转为DataFrame

完整代码

import requests
from bs4 import BeautifulSoup
import pandas as pd

# 模拟浏览器请求头,避免被反爬拦截
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
}

base_url = "https://schools.texastribune.org/districts/"
data = []

# 获取列表页内容
response = requests.get(base_url, headers=headers)
soup = BeautifulSoup(response.text, 'html.parser')

# 定位目标表格
target_table = soup.find('table', class_='table table-striped')
# 遍历表格数据行(仅处理tbody内的行,跳过表头)
for row in target_table.tbody.find_all('tr'):
    # 提取每行的详情链接
    district_link = row.find('a')
    if not district_link:
        continue
    link_href = district_link['href']
    # 拼接完整URL(处理相对链接)
    full_link = f"https://schools.texastribune.org{link_href}" if link_href.startswith('/') else link_href
    
    # 获取详情页内容
    detail_response = requests.get(full_link, headers=headers)
    detail_soup = BeautifulSoup(detail_response.text, 'html.parser')
    
    # 提取学生总数:根据页面结构定位统计区块
    total_students = None
    stats_section = detail_soup.find('div', class_='district-stats')
    if stats_section:
        for stat_item in stats_section.find_all('div', class_='stat'):
            if 'Total Students' in stat_item.get_text():
                total_students = stat_item.find('span', class_='value').get_text(strip=True)
                break
    
    # 提取地区名称
    district_name = district_link.get_text(strip=True)
    
    # 收集数据
    data.append({
        '地区名称': district_name,
        '详情链接': full_link,
        '学生总数': total_students
    })

# 转为DataFrame并输出
df = pd.DataFrame(data)
print(df)
# 可选:保存为CSV文件
# df.to_csv('德州学区学生数据.csv', index=False, encoding='utf-8-sig')

关键说明

  • 请求头设置:添加User-Agent模拟浏览器请求,降低被网站反爬机制拦截的概率
  • 精准元素定位:先找到表格主体,再遍历数据行,确保只处理有效数据
  • 详情页数据提取:根据页面的统计区块结构,精准匹配学生总数的元素
  • 容错处理:添加空值判断,避免因页面局部结构变化导致代码报错

内容的提问来源于stack exchange,提问作者Antonio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 00:44:56