使用BeautifulSoup提取表格中红绿柱状图对应数据的问题求助
Hey there! I see you're stuck trying to extract the data behind those red/green bar charts in the ShareInvestor factsheet table. Let's break down why empty1 and empty2 are coming up empty, and how to pull the actual values you need.
Why the Values Are Empty
The core issue here is that those two columns don't have plain text content. Instead, they use <div> elements to render the colored bar charts visually. When you use td.text.strip(), you're asking for text inside the <td>—but there's none there, just the bar element. That's why you're getting empty strings.
How to Extract the Bar Chart Data
To get the actual values represented by the bars, you need to look at custom data attributes on the <td> elements or the style properties of the bar divs. Here's how to adjust your code:
Step 1: Check the Element Structure (In DevTools)
First, right-click one of those bar columns and inspect it. You'll likely see something like this (varies slightly by page):
<td class="si-table__cell si-table__cell--align-center" data-value="2.8"> <div class="si-bar si-bar--negative" style="width: 56%"></div> </td>
Notice the data-value attribute? That's the actual numerical value the bar is representing. If you don't see data-value, check for other attributes like data-percent, or look at the width in the bar's style (it's usually a percentage tied to the value).
Step 2: Modify Your Code to Pull Attributes
Update your code to extract these attributes instead of just text. Here's the revised version:
import requests from bs4 import BeautifulSoup import pandas as pd # Make sure your header includes a valid User-Agent (and cookies if needed) header = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } shareInvestorRequest = requests.get( 'https://www.shareinvestor.com/fundamental/factsheet.html?counter=5259.MY', headers=header ) shareInvestorParser = BeautifulSoup(shareInvestorRequest.text, 'html.parser') try: totalShareholderReturnTable = shareInvestorParser.find('tbody', {'id':'factSheet_totalReturn'}).findAll('tr') totalShareholderReturnSelectedRow = [] for row in totalShareholderReturnTable[1:]: columns = row.findAll("td") # Extract text columns as before period = columns[0].text.strip() dividend_received = columns[1].text.strip() capital_appreciation = columns[2].text.strip() # Extract bar data from data-value attribute (adjust if your element uses a different attribute) bar1_value = columns[3].get('data-value', '') # Fallback to empty string if attribute doesn't exist bar2_value = columns[4].get('data-value', '') total_s_return = columns[5].text.strip() # Collect row data row_data = [period, dividend_received, capital_appreciation, bar1_value, bar2_value, total_s_return] totalShareholderReturnSelectedRow.append(row_data) # Debug print print(f"Extracted row: {row_data}\n") # Build your DataFrame df = pd.DataFrame( totalShareholderReturnSelectedRow, columns=['Period', 'Dividend Received', 'Capital Appreciation', 'Return Bar 1', 'Return Bar 2', 'Total Shareholder Return'] ) print("============================== TOTAL SHAREHOLDER RETURN ==============================\n") print(df) print("======================================================================================\n") except Exception as e: print(f"Oops, hit an error: {str(e)}")
Step 3: If Data Is Loaded Dynamically
If you still get empty values after this, it means the bar data is loaded via AJAX (not in the initial HTML). In that case:
- Open DevTools → Network tab → Refresh the page
- Look for XHR/fetch requests that return JSON data for the factsheet
- Copy that API URL and request it directly (you'll need to use the same headers as your browser)
- Parse the JSON to get the bar values—this is more reliable than scraping HTML
Alternatively, use selenium to simulate a browser, wait for the page to fully load, then extract the elements as above.
Quick Tips
- Always double-check the element attributes in DevTools—sites sometimes rename attributes, so
data-valuemight bedata-amountor something else - Make sure your
headerincludes a validUser-Agent; without it, ShareInvestor might block your request or return incomplete content
内容的提问来源于stack exchange,提问作者Ilham Learning

