使用BeautifulSoup4与Python提取HTML悬浮提示中的表格数字
Hey Marco, I’ve run into this exact issue with CoinGecko’s hover tooltips before—those numbers are tucked away in an HTML string inside an attribute, not directly in the page’s DOM. Let’s break down how to pull that 629925 value (and others like it) properly.
Why Your Current Code Isn’t Working
The hover details you’re after are stored in the data-original-title attribute of the .percent divs. This attribute contains a full snippet of HTML (like a table with <td> elements), but when you call developer[0].findAll("td"), BeautifulSoup is looking for <td> elements that are children of the .percent div—not the HTML string inside the attribute. That’s why you’re getting empty lists!
Step-by-Step Solution
Here’s how to extract that nested HTML data:
- Fetch and parse the main page correctly (I spotted a small typo in your code—you used
fpinstead ofresponse.text) - Pull the
data-original-titleattribute from each.percentelement - Treat the attribute’s value as a new HTML document and parse it with BeautifulSoup
- Extract the target number from this parsed tooltip HTML
Working Code Example
from bs4 import BeautifulSoup as soup import requests url = "https://www.coingecko.com/de?page=1" response = requests.get(url) # Parse the main page HTML webpage = soup(response.text, "html.parser") # Get all .percent divs containing hover tooltips developer_sections = webpage.findAll("div", {"class": "percent"}) for section in developer_sections: # Extract the tooltip's HTML string from the attribute tooltip_html = section.get("data-original-title") if tooltip_html: # Skip elements that don't have hover data # Parse the tooltip HTML as a separate soup object tooltip_soup = soup(tooltip_html, "html.parser") # Extract the target number (adjust index if your number is in a different <td>) target_number = tooltip_soup.find("td").text.strip() print(f"Extracted hover number: {target_number}")
Quick Adjustments to Note
- If your target number isn’t in the first
<td>of the tooltip table, usetooltip_soup.findAll("td")[index](replaceindexwith the position of your desired cell, e.g.,[1]for the second cell) - You can swap
html.parserwithlxmlif you have it installed—just update both soup initialization lines
This approach works because we’re treating the attribute’s HTML string as a standalone document, which lets us access the <td> elements hidden inside the hover tooltip content.
内容的提问来源于stack exchange,提问作者Marco

