You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让pandas.read_html分别读取单元格内容与tooltip而非拼接?

Solution to Split Cell Content and Tooltip in pd.read_html

Since pd.read_html concatenates visible cell text and hidden tooltip text (because the tooltip is embedded in the cell's HTML structure), the reliable fix is to manually parse the rendered HTML to extract both values separately. Here's how:

Step 1: Install Required Packages

pip install selenium beautifulsoup4 pandas html5lib

Step 2: Code to Extract Split Data

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup
import pandas as pd

# Set up headless Chrome to render dynamic content
chrome_options = Options()
chrome_options.add_argument("--headless=new")
driver = webdriver.Chrome(options=chrome_options)
driver.get("https://stats.gladiabots.com/pantheon?")

# Wait for the table to load (adjust timeout if needed)
driver.implicitly_wait(10)

# Get rendered page source
html = driver.page_source
driver.quit()

# Parse HTML with BeautifulSoup
soup = BeautifulSoup(html, "html5lib")
table = soup.find("table")

# Extract original headers
headers = [th.get_text(strip=True) for th in table.find("thead").find_all("th")]

# Create modified headers with tooltip columns for Score and XP LVL
modified_headers = []
for header in headers:
    modified_headers.append(header)
    if header in ["Score", "XP LVL"]:
        modified_headers.append(f"{header} Tooltip")

# Extract rows with split values
rows = []
for tr in table.find("tbody").find_all("tr"):
    row_cells = tr.find_all("td")
    row_data = []
    for idx, td in enumerate(row_cells):
        # Get visible text (first text node in the cell)
        visible_text = td.contents[0].strip() if td.contents else ""
        # Get tooltip text (adjust selector based on actual page HTML)
        tooltip_span = td.find("span", class_="tooltip")
        tooltip_text = tooltip_span.get_text(strip=True) if tooltip_span else td.get("title", "")
        
        row_data.append(visible_text)
        # Add tooltip column only for Score and XP LVL
        if headers[idx] in ["Score", "XP LVL"]:
            row_data.append(tooltip_text)
    rows.append(row_data)

# Create DataFrame with split columns
df = pd.DataFrame(rows, columns=modified_headers)
print(df.head())

Key Notes

  • Adjust Tooltip Selector: Inspect the page's HTML to confirm the tooltip element's class or attribute. If the tooltip is stored in a title attribute instead of a hidden span, replace the tooltip extraction line with tooltip_text = td.get("title", "").
  • Dynamic Content Handling: Selenium ensures we capture the fully rendered table, which is critical for pages that load data dynamically.
  • Faster Alternative: Check your browser's DevTools Network tab for the API endpoint that feeds the table data. Fetching JSON directly from this endpoint avoids HTML parsing entirely and is often more efficient.

内容的提问来源于stack exchange,提问作者OCa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 05:13:32