使用Beautiful Soup获取span的title属性,爬虫数据提取问题
Fixing Your Beautiful Soup Web Scraping Issue
Hey there! Let's get your transaction table scraping working correctly. The main issue with your current code is that you're pulling td elements and span elements separately—so you can't link the date from the span to the corresponding amount and transaction hash. Here's a step-by-step solution tailored to your needs:
What We'll Do
- Target the exact transaction table on the page (to avoid grabbing unrelated tables)
- Iterate over each row of the table (skipping the header)
- For each row, extract the three required fields in one go:
- The datetime from the
span'stitleattribute (first column) - The transaction amount (second column)
- The transaction hash (third column)
- The datetime from the
- Format the output exactly as you specified
Working Code
import requests from bs4 import BeautifulSoup theurl = "http://yiimp.eu/site/tx?address=DFc6oo4CAemHF4KerLG39318E1KciTs742" headers = {'User-Agent': 'MyAgent'} # Fetch the page and handle potential request errors response = requests.get(theurl, headers=headers) response.raise_for_status() # Will throw an error if the request fails (e.g., 404/500) soup = BeautifulSoup(response.text, "html.parser") # Locate the transaction table using its CSS classes (matches the page's structure) transaction_table = soup.find('table', class_='table table-striped table-hover') if transaction_table: # Skip the header row (first row in the table) and loop through data rows data_rows = transaction_table.find_all('tr')[1:] for row in data_rows: # Get all <td> elements in the current row columns = row.find_all('td') # Ensure we have at least 3 columns to avoid index errors if len(columns) >= 3: # Extract datetime from the span's title attribute transaction_datetime = columns[0].find('span')['title'] # Extract amount (strip removes extra whitespace/newlines) transaction_amount = columns[1].get_text(strip=True) # Extract transaction hash transaction_hash = columns[2].get_text(strip=True) # Print in your desired format print(f"{transaction_datetime} {transaction_amount} {transaction_hash}") else: print("Oops, couldn't find the transaction table on the page.")
Key Improvements Over Your Original Code
- Contextual Extraction: Instead of pulling all
tds andspans globally, we extract data row by row—so each datetime is automatically paired with its correct amount and hash. - Error Handling:
response.raise_for_status()ensures we catch issues like broken links or server errors early. - Clean Text:
get_text(strip=True)removes unwanted spaces and line breaks from the amount and hash. - Precise Table Targeting: Using the table's CSS classes ensures we don't accidentally scrape other tables on the page.
Why Your Original Code Had Issues
- Separating the
tdandspanloops meant you couldn't associate each date with the right transaction data. - The
countlogic was broken (count ==0is a comparison, not an assignment, and thebreakwas misplaced), so it didn't help with duplication.
内容的提问来源于stack exchange,提问作者Jcmoney1010
相关产品推荐
相关产品推荐

