You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium如何提取网页指定标题下所有链接内的表格?

Got it, let's break down how to solve this problem—extracting all tables from links under a specific heading on a webpage. We'll use Python with standard web scraping tools, which is the most straightforward approach here.

Step 1: Install Required Libraries

First, make sure you have these packages installed (they handle fetching pages, parsing HTML, and working with tables):

pip install requests beautifulsoup4 pandas
Step 2: Full Code Implementation

Here's a complete, commented script that does exactly what you need. I've included error handling and flexibility for different page structures:

import requests
from bs4 import BeautifulSoup
import pandas as pd
from urllib.parse import urljoin

def extract_tables_under_title(base_url, target_title):
    # Fetch the main page with the target heading
    try:
        response = requests.get(base_url)
        response.raise_for_status()  # Throw error if page doesn't load properly
    except requests.exceptions.RequestException as e:
        print(f"Failed to load main page: {e}")
        return []

    soup = BeautifulSoup(response.text, 'html.parser')

    # Locate the target heading (adjust tag list if your title is h1/h3/h4 instead of h2)
    target_heading = soup.find(
        lambda tag: tag.name in ['h1', 'h2', 'h3', 'h4'] 
        and target_title.strip().lower() in tag.get_text(strip=True).lower()
    )
    if not target_heading:
        print(f"Couldn't find the heading: '{target_title}'")
        return []

    # Extract all links under the heading (stops at the next heading to avoid grabbing unrelated links)
    links = []
    current_element = target_heading.next_sibling
    
    while current_element and current_element.name not in ['h1', 'h2', 'h3', 'h4']:
        # Check if the current element is a link itself
        if current_element.name == 'a' and current_element.get('href'):
            absolute_link = urljoin(base_url, current_element['href'])
            links.append(absolute_link)
        # Check for links inside nested elements (like paragraphs or divs)
        if hasattr(current_element, 'find_all'):
            for a_tag in current_element.find_all('a', href=True):
                absolute_link = urljoin(base_url, a_tag['href'])
                if absolute_link not in links:
                    links.append(absolute_link)
        current_element = current_element.next_sibling

    if not links:
        print(f"No links found under the heading: '{target_title}'")
        return []

    # Extract tables from each linked page
    all_tables = []
    for idx, link in enumerate(links, 1):
        print(f"Processing link {idx}/{len(links)}: {link}")
        try:
            link_response = requests.get(link)
            link_response.raise_for_status()
            # Pandas automatically finds and parses all HTML tables on the page
            page_tables = pd.read_html(link_response.text)
            if page_tables:
                all_tables.extend(page_tables)
                print(f"Pulled {len(page_tables)} table(s) from this link")
            else:
                print(f"No tables found on this page")
        except requests.exceptions.RequestException as e:
            print(f"Failed to load link: {e}")
        except ValueError as e:
            print(f"Couldn't parse tables from this link: {e}")

    return all_tables

# Example usage (replace these values with your actual URL and heading)
if __name__ == "__main__":
    MAIN_PAGE_URL = "https://your-target-page.com"
    TARGET_HEADING_TEXT = "Your Specified Heading Here"
    
    extracted_tables = extract_tables_under_title(MAIN_PAGE_URL, TARGET_HEADING_TEXT)
    
    # Save each table to a separate CSV file
    for table_num, table in enumerate(extracted_tables, 1):
        table.to_csv(f"extracted_table_{table_num}.csv", index=False)
        print(f"Saved table {table_num} to extracted_table_{table_num}.csv")
Step 3: Customize for Your Page

You'll need to tweak a few parts based on your specific webpage:

  • Heading tag: If your target title is an <h3> instead of <h2>, update the tag list in the target_heading line.
  • Link container: If all links are inside a specific div (e.g., <div class="resource-links"> right after the heading), replace the link-traversal loop with something simpler:
    # Example: If links are in a dedicated container after the heading
    link_container = target_heading.find_next_sibling('div', class_='resource-links')
    if link_container:
        links = [urljoin(base_url, a['href']) for a in link_container.find_all('a', href=True)]
    
  • Dynamic content: If the page loads links via JavaScript (e.g., React/Vue apps), swap requests with selenium to render the page first—just let me know if you need help with that!

内容的提问来源于stack exchange,提问作者user9764822

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:37:44