You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Beautiful Soup爬取Pregame网页数据失败求助

问题:使用BeautifulSoup爬取Pregame网页无法获取完整表格数据

爬取网页https://pregame.com/game-center/171763/consensus-archive时,仅能获取少量HTML片段,无法提取网页中可见的表格嵌入数据。原本目标是抓取整个表格,目前尝试抓取日期列也未成功。

现有代码

import pandas as pd
import requests
from bs4 import BeautifulSoup
url = 'https://pregame.com/game-center/171763/consensus-archive'
html = requests.get(url)
soup = BeautifulSoup(html.text, 'html.parser')
results = soup.find(class_ = "pg-move-list")
print(results.prettify())
dates = results.find_all("td", class_="pg-col pg-col--date")
for date in dates:
    print(date, end = "\n"*2)
for date in dates:
    data_date = date.find("p", class_= "pg-col-data")
print(data_date.text)

网页HTML结构截图

网页HTML结构截图


解决方案

问题原因

该网站的表格数据是通过JavaScript动态渲染的,requests.get()只能获取页面初始加载的静态HTML,无法获取JS执行后生成的内容,所以你拿到的只是空的或不完整的表格框架。

方法1:使用Selenium模拟浏览器渲染

通过Selenium启动真实浏览器加载页面,等待JS渲染完成后再获取完整页面源码,即可提取表格数据。

代码示例

import pandas as pd
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

url = 'https://pregame.com/game-center/171763/consensus-archive'

# 初始化Chrome浏览器(需提前下载对应版本的chromedriver并配置环境变量)
driver = webdriver.Chrome()
driver.get(url)

# 等待表格元素加载完成,超时时间10秒
wait = WebDriverWait(driver, 10)
wait.until(EC.presence_of_element_located((By.CLASS_NAME, "pg-move-list")))

# 获取渲染后的完整页面源码
html = driver.page_source
soup = BeautifulSoup(html, 'html.parser')

# 提取整个表格数据
table = soup.find(class_="pg-move-list")
rows = table.find_all("tr")

# 解析表头和行数据
headers = [th.get_text(strip=True) for th in rows[0].find_all("th")]
table_data = []
for row in rows[1:]:
    columns = [td.get_text(strip=True) for td in row.find_all("td")]
    table_data.append(columns)

# 转成DataFrame方便后续处理
df = pd.DataFrame(table_data, columns=headers)
print(df)

# 关闭浏览器
driver.quit()

方法2:直接请求数据接口(更高效)

打开浏览器开发者工具(F12),切换到Network标签,刷新页面后筛选XHR/Fetch请求,查找返回表格数据的API接口。找到后直接用requests请求该接口,解析JSON数据即可,无需模拟浏览器。

示例代码(需替换为实际找到的接口URL):

import pandas as pd
import requests

api_url = "实际找到的API接口URL"
response = requests.get(api_url)
data = response.json()

# 根据接口返回的JSON结构解析数据,转成DataFrame
df = pd.DataFrame(data)
print(df)

内容的提问来源于stack exchange,提问作者Abb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 10:15:33