You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup解析表格仅获部分内容的解决求助

问题原因

backpack.tf的表格采用动态加载机制:初始页面仅渲染约100条数据,剩余条目会在用户滚动页面时通过AJAX异步请求加载。requests库只能获取页面首次加载的静态HTML,自然无法拿到完整的2000+条数据,更换解析器也解决不了这个问题——因为问题根源不在解析器,而是数据未被包含在初始响应里。

解决方案

方案1:用Selenium模拟浏览器加载全部数据

Selenium可以模拟真实浏览器的行为,滚动页面触发数据加载,等所有内容渲染完成后再解析:

  1. 先安装依赖:
pip install selenium
  1. 下载对应浏览器的驱动(比如Chrome的ChromeDriver),确保驱动版本和浏览器版本匹配,将驱动路径配置到环境变量或脚本中。

  2. 示例代码:

from bs4 import BeautifulSoup
from selenium import webdriver
import time

url = "https://backpack.tf/spreadsheet"

# 初始化Chrome浏览器(无头模式,不显示窗口)
options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
driver = webdriver.Chrome(options=options)
driver.get(url)

# 滚动页面直到加载完所有数据
last_height = driver.execute_script("return document.body.scrollHeight")
while True:
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(2)  # 等待数据加载
    new_height = driver.execute_script("return document.body.scrollHeight")
    if new_height == last_height:
        break
    last_height = new_height

# 解析完整页面内容
soup = BeautifulSoup(driver.page_source, "html.parser")
for tr in soup.find_all("tr"):
    print(tr.text.strip())

driver.quit()

方案2:直接调用网站的数据API(更高效)

打开浏览器开发者工具(F12),切换到「网络」标签,滚动页面时观察XHR请求,找到加载表格数据的API接口,直接请求该接口获取JSON格式的完整数据,无需解析HTML:

示例(需自行确认实际API接口):

import requests

# 替换为实际找到的API接口URL
api_url = "https://backpack.tf/api/spreadsheet/data"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

response = requests.get(api_url, headers=headers)
data = response.json()

# 处理JSON数据,示例:打印条目信息
for item in data.get("items", []):
    print(item.get("name", "未知名称"))

注意:API接口可能存在反爬限制,需要携带正确的请求头,部分接口可能需要登录状态的Cookie,需根据实际情况调整。

内容的提问来源于stack exchange,提问作者maxim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 21:36:26