如何用BeautifulSoup提取网页指定区间内的.mid格式链接
Got it, let's solve this problem so you can grab those Reels-specific .mid links easily, and reuse the code for other similar sites later.
The Core Idea
First, we need to pinpoint the exact section of the page bounded by your start/end identifiers (like the "Reels" heading and the next major heading that marks the end of the Reels section). Then, we'll only extract .mid links within that specific block.
Reusable Python Code
Here's a flexible function that works for this scenario, with comments to explain each step:
import requests from bs4 import BeautifulSoup def extract_section_links(url, start_heading_text, end_heading_text, target_extension): # Fetch the page content and handle HTTP errors response = requests.get(url) response.raise_for_status() soup = BeautifulSoup(response.text, "html.parser") # Locate the start of the target section (e.g., the "Reels" heading) start_heading = soup.find(lambda tag: tag.name in ["h2", "h3", "h4"] and start_heading_text in tag.get_text(strip=True)) if not start_heading: print(f"Error: Couldn't find start heading with text '{start_heading_text}'") return [] # Traverse elements from the start heading until we hit the end heading current_element = start_heading.next_sibling target_links = [] while current_element: # Check if we've reached the end of the section if current_element.name in ["h2", "h3", "h4"] and end_heading_text in current_element.get_text(strip=True): break # Extract all matching links from the current element and its children for link in current_element.find_all("a", href=True): link_href = link["href"] if link_href.endswith(target_extension): # Convert relative URLs to full, valid URLs full_link = requests.compat.urljoin(url, link_href) target_links.append(full_link) # Move to the next element in the page structure current_element = current_element.next_sibling return target_links
How to Use It for Your Target Site
Looking at the source of your target page, the Reels section starts with <h3>Reels</h3> and ends right before <h3>Jigs</h3>. Call the function like this:
# Example usage for tadpoletunes.com page_url = "http://www.tadpoletunes.com/tunes/celtic1/" reels_mid_links = extract_section_links( url=page_url, start_heading_text="Reels", end_heading_text="Jigs", target_extension=".mid" ) # Print the collected links print(f"Found {len(reels_mid_links)} Reels .mid links:") for link in reels_mid_links: print(link)
Key Tips for Reusing on Other Sites
- Adjust heading tags: If the site uses
<h2>instead of<h3>for section headers, update thetag.name in ["h2", "h3", "h4"]part to match. - Update start/end text: Replace "Reels" and "Jigs" with the actual section headers from the target site (check the page source to find these boundary markers).
- Switch extensions: If you need links other than .mid, just swap
.midwith your target file extension (like.mp3or.abc).
Pro tip: To find the right start/end identifiers, view the page source and look for the text that immediately precedes/follows the section you care about. It's almost always a heading tag (h2-h4) or a bolded section title.
内容的提问来源于stack exchange,提问作者steveeweeveewoo

