You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取网页指定区间内的.mid格式链接

Got it, let's solve this problem so you can grab those Reels-specific .mid links easily, and reuse the code for other similar sites later.

The Core Idea

First, we need to pinpoint the exact section of the page bounded by your start/end identifiers (like the "Reels" heading and the next major heading that marks the end of the Reels section). Then, we'll only extract .mid links within that specific block.

Reusable Python Code

Here's a flexible function that works for this scenario, with comments to explain each step:

import requests
from bs4 import BeautifulSoup

def extract_section_links(url, start_heading_text, end_heading_text, target_extension):
    # Fetch the page content and handle HTTP errors
    response = requests.get(url)
    response.raise_for_status()
    soup = BeautifulSoup(response.text, "html.parser")

    # Locate the start of the target section (e.g., the "Reels" heading)
    start_heading = soup.find(lambda tag: tag.name in ["h2", "h3", "h4"] and start_heading_text in tag.get_text(strip=True))
    if not start_heading:
        print(f"Error: Couldn't find start heading with text '{start_heading_text}'")
        return []

    # Traverse elements from the start heading until we hit the end heading
    current_element = start_heading.next_sibling
    target_links = []

    while current_element:
        # Check if we've reached the end of the section
        if current_element.name in ["h2", "h3", "h4"] and end_heading_text in current_element.get_text(strip=True):
            break

        # Extract all matching links from the current element and its children
        for link in current_element.find_all("a", href=True):
            link_href = link["href"]
            if link_href.endswith(target_extension):
                # Convert relative URLs to full, valid URLs
                full_link = requests.compat.urljoin(url, link_href)
                target_links.append(full_link)

        # Move to the next element in the page structure
        current_element = current_element.next_sibling

    return target_links

How to Use It for Your Target Site

Looking at the source of your target page, the Reels section starts with <h3>Reels</h3> and ends right before <h3>Jigs</h3>. Call the function like this:

# Example usage for tadpoletunes.com
page_url = "http://www.tadpoletunes.com/tunes/celtic1/"
reels_mid_links = extract_section_links(
    url=page_url,
    start_heading_text="Reels",
    end_heading_text="Jigs",
    target_extension=".mid"
)

# Print the collected links
print(f"Found {len(reels_mid_links)} Reels .mid links:")
for link in reels_mid_links:
    print(link)

Key Tips for Reusing on Other Sites

  1. Adjust heading tags: If the site uses <h2> instead of <h3> for section headers, update the tag.name in ["h2", "h3", "h4"] part to match.
  2. Update start/end text: Replace "Reels" and "Jigs" with the actual section headers from the target site (check the page source to find these boundary markers).
  3. Switch extensions: If you need links other than .mid, just swap .mid with your target file extension (like .mp3 or .abc).

Pro tip: To find the right start/end identifiers, view the page source and look for the text that immediately precedes/follows the section you care about. It's almost always a heading tag (h2-h4) or a bolded section title.

内容的提问来源于stack exchange,提问作者steveeweeveewoo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:28:59