You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Beautiful Soup爬取网页表格时DataFrame为空的问题求助

问题描述

尝试使用Beautiful Soup爬取指定页面的表格,但始终得到空的DataFrame,目标是生成包含Metro Area、Median Home Price和Affordability Index三列的DataFrame。原实现代码如下:

import requests
import pandas as pd
from bs4 import BeautifulSoup

URL = "https://www.kiplinger.com/article/real-estate/t010-c000-s002-home-price-changes-in-the-100-largest-metro-areas.html"
page = requests.get(URL)

soup = BeautifulSoup(page.content, 'html.parser')
table = soup.find('table')

data = []
for tr in table.find_all('tr'):
 row = {}
 cells = tr.find_all('td')
  if len(cells) == 3:
    row['Metro Area'] = cells[0].text.strip()
    row['Median Home Price'] = cells[1].text.strip()
    row['Affordability Index'] = cells[2].text.strip()
    data.append(row)


df = pd.DataFrame(data)
print(df)

问题原因及修复方案

核心问题

  1. 缩进错误:原代码中row = {}和if语句的缩进不符合Python语法规范,导致逻辑无法正常执行。
  2. 表格定位不准确:直接用soup.find('table')可能无法精准匹配目标表格,页面中的表格带有特定类名,需要更精确的选择器。
  3. 反爬限制:网站可能拦截无请求头的爬虫请求,返回空内容或非目标页面。

修正后的代码

import requests
import pandas as pd
from bs4 import BeautifulSoup

URL = "https://www.kiplinger.com/article/real-estate/t010-c000-s002-home-price-changes-in-the-100-largest-metro-areas.html"
# 添加请求头模拟浏览器访问
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}
page = requests.get(URL, headers=headers)

soup = BeautifulSoup(page.content, 'html.parser')
# 通过class精准定位目标表格
table = soup.find('table', class_="table")

data = []
# 修正缩进,确保循环逻辑正确
for tr in table.find_all('tr'):
    row = {}
    cells = tr.find_all('td')
    # 只处理包含3个数据单元格的行(自动排除表头行)
    if len(cells) == 3:
        row['Metro Area'] = cells[0].text.strip()
        row['Median Home Price'] = cells[1].text.strip()
        row['Affordability Index'] = cells[2].text.strip()
        data.append(row)

df = pd.DataFrame(data)
print(df)

关键修复点说明

  • 请求头配置:添加User-Agent字段,避免被网站识别为爬虫,确保能获取到完整页面内容。
  • 表格精准选择:使用class_="table"匹配页面中的目标表格,避免定位到无关表格元素。
  • 缩进修正:调整代码缩进,符合Python语法要求,保证循环和条件判断逻辑正常执行。
  • 自动排除表头:表头行使用<th>标签而非<td>,通过判断<td>数量自然过滤表头,无需额外处理。

内容的提问来源于stack exchange,提问作者user20901436

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 07:01:03