如何用Python抓取JavaScript饼图的节点文本?含Dice站点示例
Hey there! Let's tackle these two questions about scraping JavaScript pie charts with Python—this is a common pain point, but we've got clear solutions for both the general case and the specific Dice page.
Question 1: General Method to Grab Text from JS Pie Charts
JavaScript pie charts typically store their data in one of two ways, so we’ll cover both scenarios to make sure you’re covered:
1. Extract Data Directly from the DOM (No Browser Rendering Required)
Lots of chart libraries (like Highcharts, Chart.js) don’t just render the chart—they also dump the raw data into the page’s DOM, usually in a hidden <script> tag or a data attribute. This is the easiest case because you don’t need to fire up a browser:
- Use
requeststo fetch the page’s HTML - Parse it with BeautifulSoup to find the script tag holding the chart data
- Pull out the JSON string, parse it, and extract your labels/text
Here’s a quick example:
import requests from bs4 import BeautifulSoup import json url = "YOUR_TARGET_PAGE_URL" response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") # Look for a script tag that mentions your chart data (adjust the keyword to match your use case) chart_script = soup.find("script", text=lambda t: t and "pieChartData" in t) if chart_script: # Clean up the string to get valid JSON (you might need to tweak this based on the page) data_str = chart_script.text.split("var pieChartData =")[1].split(";")[0].strip() pie_data = json.loads(data_str) # Grab the labels/text nodes from the parsed data pie_labels = [item["label"] for item in pie_data] print("Pie chart text nodes:", pie_labels)
2. Render the Page with a Headless Browser (For Dynamic Canvas/SVG Charts)
If the chart is drawn directly to a <canvas> or <svg> without exposing the raw data in the DOM, you’ll need to render the page to access the visible text nodes. Tools like Selenium or Playwright are perfect for this—they simulate a real browser and let you interact with the rendered page.
Selenium Example:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC url = "YOUR_TARGET_PAGE_URL" # Initialize Chrome (make sure you have ChromeDriver installed and in your PATH) driver = webdriver.Chrome() driver.get(url) # Wait for the pie chart to fully load (adjust the selector to match your chart's elements) wait = WebDriverWait(driver, 10) # This selector targets text elements inside the pie chart's data labels (adjust as needed) pie_text_nodes = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "svg .highcharts-data-label text"))) # Extract and clean the text from each node labels = [node.text.strip() for node in pie_text_nodes] print("Pie chart text nodes:", labels) # Don't forget to close the browser driver.quit()
Question 2: Scrape the JavaScript Pie Chart on https://www.dice.com/skills/javascript
Now let’s get specific to the Dice skills page. After checking the page, the skill distribution pie chart uses Highcharts, which renders labels as SVG text elements. We can either scrape the rendered text or pull the raw data directly from the DOM—both work great.
Option 1: Scrape Rendered Text with Selenium
This grabs the visible text nodes directly from the rendered pie chart:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC url = "https://www.dice.com/skills/javascript" driver = webdriver.Chrome() driver.get(url) # Wait for the pie chart's data labels to load (targets the tspan elements inside SVG text) wait = WebDriverWait(driver, 15) skill_labels = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "g.highcharts-data-label text tspan"))) # Extract only non-empty text nodes pie_node_texts = [label.text.strip() for label in skill_labels if label.text.strip()] print("JavaScript pie chart node texts:", pie_node_texts) driver.quit()
Option 2: Extract Raw Data from the DOM (No Browser Needed)
If you want to avoid using a browser, the Dice page stores the raw skill data in a <script> tag. We can parse that directly:
import requests from bs4 import BeautifulSoup import json url = "https://www.dice.com/skills/javascript" response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") # Find the script tag containing the skill distribution data skill_script = soup.find("script", text=lambda t: t and "skillDistribution" in t) if skill_script: # Clean the string to extract valid JSON data_start = skill_script.text.find('"skillDistribution":') + len('"skillDistribution":') data_end = skill_script.text.find('}', data_start) + 1 skill_data = json.loads(skill_script.text[data_start:data_end]) # The "name" field in each skill entry is the pie chart node text skill_names = [item["name"] for item in skill_data["skills"]] print("Pie chart node texts:", skill_names)
内容的提问来源于stack exchange,提问作者Ali osama

