Python新手求助:如何创建可检测特定按钮内链接变化的网站更新工具
Hey there! Since you're new to Python and trying to build a tool to monitor a specific button's link on that Kelda Group page, let's break this down into easy, actionable steps.
Your current code grabs every <a> tag on the page, which is why you're seeing so many links. Since your target link is in a fixed position, we can use HTML element attributes (like class, ID, or surrounding parent elements) to pinpoint it directly.
How to find the right attributes:
Right-click the target button on the webpage → select "Inspect" to view its HTML structure. Look for unique identifiers like:
- A specific
class(e.g.,btn-creditor-document) - A unique
id(e.g.,latest-creditor-link) - Exact button text (e.g., "Download Creditor Report")
- A parent container with a distinct class/id (e.g., a
<div>wrapping the button)
Example code for precise targeting:
import requests from bs4 import BeautifulSoup url = 'https://www.keldagroup.com/investors/creditor-considerations/' reqs = requests.get(url) soup = BeautifulSoup(reqs.text, 'html.parser') # Option 1: Target by button class (replace with your button's actual class) target_link = soup.find('a', class_='btn-creditor-document') # Option 2: Target by button text (match the exact text on the button) # target_link = soup.find('a', string='Download Creditor Report') # Option 3: Target via parent container (e.g., a div with class "creditor-section") # target_link = soup.select_one('div.creditor-section a:first-of-type') # Extract and print the link if found if target_link: href = target_link.get('href') print(f"Target link found: {href}") else: print("Could not locate the target link—double-check your selector!")
Once you can reliably fetch the target link, you can compare it to a saved version to check for updates. Here's a simple implementation:
import requests from bs4 import BeautifulSoup def get_target_link(): url = 'https://www.keldagroup.com/investors/creditor-considerations/' reqs = requests.get(url) soup = BeautifulSoup(reqs.text, 'html.parser') # Use your confirmed selector here target_link = soup.select_one('div.creditor-section a.btn') return target_link.get('href') if target_link else None def check_link_update(): # Read the last saved link (if exists) try: with open('saved_link.txt', 'r') as f: old_link = f.read().strip() except FileNotFoundError: old_link = None new_link = get_target_link() if not new_link: print("Error: Failed to fetch target link.") return if not old_link: print("First run—saving initial link.") with open('saved_link.txt', 'w') as f: f.write(new_link) elif new_link != old_link: print(f"Link updated! Old: {old_link}\nNew: {new_link}") # Add notifications here (e.g., email, desktop alert) if needed with open('saved_link.txt', 'w') as f: f.write(new_link) else: print("No changes detected.") if __name__ == "__main__": check_link_update()
For a more robust setup with less manual code, consider these tools:
requests-cache: Caches webpage requests to avoid redundant downloads (saves bandwidth and speeds up checks). Install withpip install requests-cache.schedule: Automates periodic checks (e.g., run daily at 9 AM). Install withpip install schedule.
Example with both tools:
import requests from bs4 import BeautifulSoup import requests_cache import schedule import time # Cache requests for 1 hour to avoid hitting the site too often requests_cache.install_cache('kelda_cache', expire_after=3600) def get_target_link(): url = 'https://www.keldagroup.com/investors/creditor-considerations/' reqs = requests.get(url) soup = BeautifulSoup(reqs.text, 'html.parser') return soup.select_one('div.creditor-section a.btn').get('href') def check_link_update(): try: with open('saved_link.txt', 'r') as f: old_link = f.read().strip() except FileNotFoundError: old_link = None new_link = get_target_link() if new_link != old_link: print(f"Link updated! Old: {old_link}\nNew: {new_link}") with open('saved_link.txt', 'w') as f: f.write(new_link) else: print("No changes.") # Schedule daily checks at 9 AM schedule.every().day.at("09:00").do(check_link_update) print("Monitoring link changes... Press Ctrl+C to stop.") while True: schedule.run_pending() time.sleep(60)
内容的提问来源于stack exchange,提问作者van10

