Python2.7+BeautifulSoup4提取加密货币涨幅数据的技术问题
Hey there! Let’s sort this out for you—you don’t need to write the <tr> content to a file first; you can parse the HTML directly in memory with BeautifulSoup4, which is way more efficient. Here’s a step-by-step solution tailored for Python 2.7:
Step-by-Step Approach
- First, fetch the page HTML (using
urllib2since we’re on Python 2.7) - Parse the HTML with BeautifulSoup
- Target the
<tr>tags by theiridpattern - Extract the coin name and time period from the
idstring - Grab the corresponding percentage change from the nested
<td>tag - Package everything into a list of dictionaries
Working Code Example
from bs4 import BeautifulSoup import urllib2 # Fetch the gainers/losers page (update the URL if needed) url = "https://coinmarketcap.com/gainers-losers/" try: response = urllib2.urlopen(url) html_content = response.read() except urllib2.URLError as e: print("Error fetching page: {}".format(e)) exit() # Parse the HTML with BeautifulSoup (use 'lxml' instead if you have it installed) soup = BeautifulSoup(html_content, 'html.parser') # Initialize an empty list to store our results coin_changes = [] # Find all <tr> tags where the ID starts with "id-gainers-" for tr in soup.find_all('tr', id=lambda x: x and x.startswith("id-gainers-")): # Split the ID string to extract coin name and time period id_components = tr['id'].split("-") coin_name = id_components[2] time_frame = id_components[3] # Find the <td> tag that matches the time frame (handles positive/negative changes) change_td = tr.find('td', class_=lambda x: x and "percent-{}".format(time_frame) in x) if change_td: percentage = change_td.get_text(strip=True) # Add the data to our list as a dictionary coin_changes.append({ "coin_name": coin_name, "time_period": time_frame, "change_percent": percentage }) # Print the final result for entry in coin_changes: print("Coin: {}, {} Change: {}".format(entry['coin_name'], entry['time_period'], entry['change_percent']))
Key Notes
- No file writing needed: Parsing directly from the fetched HTML avoids unnecessary disk I/O and simplifies your code.
- Flexible selectors: Using
lambdaforidandclassmatching lets you handle both gainers and losers (just swap"id-gainers-"with"id-losers-"for the latter) and works for positive/negative change classes. - String splitting: The
idformat (id-gainers-adtoken-1h) is consistent, so splitting on-gives us reliable access to the coin name and time period.
If you already have the HTML content saved locally (instead of fetching it live), you can skip the urllib2 part and just pass the saved HTML string directly to BeautifulSoup.
内容的提问来源于stack exchange,提问作者Forum Studiologe
相关产品推荐
相关产品推荐

