You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用requests与beautifulsoup4获取tbody数据时标签为空的问题求助

这个问题我太熟了!你碰到的是动态网页加载的典型坑——用requests拿到的只是页面的「骨架」,而浏览器里看到的表格数据是JavaScript后来渲染出来的,所以BeautifulSoup解析静态HTML时,<tbody>自然是空的。给你两个靠谱的解决方案:

方案一:用Selenium模拟浏览器渲染页面

Selenium会像真实用户一样打开浏览器、加载页面,能获取到JS渲染后的完整HTML内容。步骤如下:

  1. 先安装依赖:
pip install selenium
  1. 下载对应浏览器的驱动(比如Chrome的ChromeDriver),确保驱动版本和你的浏览器版本匹配,然后记住驱动的路径。
  2. 编写代码:
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from bs4 import BeautifulSoup
import time

def get_riven_data(mod_name):
    # 构造目标URL,注意把mod名转成小写
    url = f'https://semlar.com/rivenprices/{mod_name.lower()}'
    
    # 初始化Chrome浏览器(替换成你的ChromeDriver绝对路径)
    service = Service('你的ChromeDriver路径')
    driver = webdriver.Chrome(service=service)
    
    # 打开页面并等待加载完成(时间可根据网络情况调整)
    driver.get(url)
    time.sleep(3)
    
    # 获取渲染后的完整页面源码
    page_source = driver.page_source
    driver.quit()  # 关闭浏览器
    
    # 用BeautifulSoup解析
    soup = BeautifulSoup(page_source, 'html.parser')
    table = soup.find(class_='table')
    
    if not table:
        print(f"没找到关于{mod_name}的数据哦")
        return
    
    # 提取tbody里的内容
    tbody = table.find('tbody')
    rows = tbody.find_all('tr')
    for row in rows:
        cols = row.find_all('td')
        if len(cols) >= 3:
            avg_price = cols[1].text.strip()
            dispo = cols[2].text.strip()
            print(f"Avg Price: {avg_price}, Dispo: {dispo}")

# 接收用户输入并调用函数
mod_name = input("请输入Mod名称(比如Artax、Lanka):")
get_riven_data(mod_name)
方案二:直接调用网站的API接口(推荐)

其实这个网站是通过API接口获取数据的,直接请求API比模拟浏览器高效太多,还不用装驱动。它的API格式很简单:

请求URL是 https://semlar.com/api/rivenprices?item=Mod名称,返回的是JSON格式的数据,直接解析就行。代码示例:

import requests

def get_riven_data(mod_name):
    api_url = f'https://semlar.com/api/rivenprices?item={mod_name.lower()}'
    response = requests.get(api_url)
    
    if response.status_code != 200:
        print(f"请求出错啦,状态码:{response.status_code}")
        return
    
    data = response.json()
    if not data:
        print(f"没找到关于{mod_name}的数据哦")
        return
    
    # 遍历数据并输出
    for item in data:
        avg_price = item.get('avg_price')
        dispo = item.get('dispo')
        print(f"Avg Price: {avg_price}, Dispo: {dispo}")

# 接收用户输入
mod_name = input("请输入Mod名称(比如Artax、Lanka):")
get_riven_data(mod_name)

两种方案对比

  • API方案:速度快、资源占用少,是最优选择,只要能找到接口就优先用它。
  • Selenium方案:适合那些找不到API的复杂动态页面,但需要额外安装驱动,运行速度慢一些。

内容的提问来源于stack exchange,提问作者Sqoshu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:39:09