如何让pandas.read_html分别读取单元格内容与tooltip而非拼接?
Solution to Split Cell Content and Tooltip in pd.read_html
Since pd.read_html concatenates visible cell text and hidden tooltip text (because the tooltip is embedded in the cell's HTML structure), the reliable fix is to manually parse the rendered HTML to extract both values separately. Here's how:
Step 1: Install Required Packages
pip install selenium beautifulsoup4 pandas html5lib
Step 2: Code to Extract Split Data
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup import pandas as pd # Set up headless Chrome to render dynamic content chrome_options = Options() chrome_options.add_argument("--headless=new") driver = webdriver.Chrome(options=chrome_options) driver.get("https://stats.gladiabots.com/pantheon?") # Wait for the table to load (adjust timeout if needed) driver.implicitly_wait(10) # Get rendered page source html = driver.page_source driver.quit() # Parse HTML with BeautifulSoup soup = BeautifulSoup(html, "html5lib") table = soup.find("table") # Extract original headers headers = [th.get_text(strip=True) for th in table.find("thead").find_all("th")] # Create modified headers with tooltip columns for Score and XP LVL modified_headers = [] for header in headers: modified_headers.append(header) if header in ["Score", "XP LVL"]: modified_headers.append(f"{header} Tooltip") # Extract rows with split values rows = [] for tr in table.find("tbody").find_all("tr"): row_cells = tr.find_all("td") row_data = [] for idx, td in enumerate(row_cells): # Get visible text (first text node in the cell) visible_text = td.contents[0].strip() if td.contents else "" # Get tooltip text (adjust selector based on actual page HTML) tooltip_span = td.find("span", class_="tooltip") tooltip_text = tooltip_span.get_text(strip=True) if tooltip_span else td.get("title", "") row_data.append(visible_text) # Add tooltip column only for Score and XP LVL if headers[idx] in ["Score", "XP LVL"]: row_data.append(tooltip_text) rows.append(row_data) # Create DataFrame with split columns df = pd.DataFrame(rows, columns=modified_headers) print(df.head())
Key Notes
- Adjust Tooltip Selector: Inspect the page's HTML to confirm the tooltip element's class or attribute. If the tooltip is stored in a
titleattribute instead of a hidden span, replace the tooltip extraction line withtooltip_text = td.get("title", ""). - Dynamic Content Handling: Selenium ensures we capture the fully rendered table, which is critical for pages that load data dynamically.
- Faster Alternative: Check your browser's DevTools Network tab for the API endpoint that feeds the table data. Fetching JSON directly from this endpoint avoids HTML parsing entirely and is often more efficient.
内容的提问来源于stack exchange,提问作者OCa
相关产品推荐
相关产品推荐

