Selenium如何提取网页指定标题下所有链接内的表格?
Got it, let's break down how to solve this problem—extracting all tables from links under a specific heading on a webpage. We'll use Python with standard web scraping tools, which is the most straightforward approach here.
Step 1: Install Required Libraries
First, make sure you have these packages installed (they handle fetching pages, parsing HTML, and working with tables):
pip install requests beautifulsoup4 pandas
Step 2: Full Code Implementation
Here's a complete, commented script that does exactly what you need. I've included error handling and flexibility for different page structures:
import requests from bs4 import BeautifulSoup import pandas as pd from urllib.parse import urljoin def extract_tables_under_title(base_url, target_title): # Fetch the main page with the target heading try: response = requests.get(base_url) response.raise_for_status() # Throw error if page doesn't load properly except requests.exceptions.RequestException as e: print(f"Failed to load main page: {e}") return [] soup = BeautifulSoup(response.text, 'html.parser') # Locate the target heading (adjust tag list if your title is h1/h3/h4 instead of h2) target_heading = soup.find( lambda tag: tag.name in ['h1', 'h2', 'h3', 'h4'] and target_title.strip().lower() in tag.get_text(strip=True).lower() ) if not target_heading: print(f"Couldn't find the heading: '{target_title}'") return [] # Extract all links under the heading (stops at the next heading to avoid grabbing unrelated links) links = [] current_element = target_heading.next_sibling while current_element and current_element.name not in ['h1', 'h2', 'h3', 'h4']: # Check if the current element is a link itself if current_element.name == 'a' and current_element.get('href'): absolute_link = urljoin(base_url, current_element['href']) links.append(absolute_link) # Check for links inside nested elements (like paragraphs or divs) if hasattr(current_element, 'find_all'): for a_tag in current_element.find_all('a', href=True): absolute_link = urljoin(base_url, a_tag['href']) if absolute_link not in links: links.append(absolute_link) current_element = current_element.next_sibling if not links: print(f"No links found under the heading: '{target_title}'") return [] # Extract tables from each linked page all_tables = [] for idx, link in enumerate(links, 1): print(f"Processing link {idx}/{len(links)}: {link}") try: link_response = requests.get(link) link_response.raise_for_status() # Pandas automatically finds and parses all HTML tables on the page page_tables = pd.read_html(link_response.text) if page_tables: all_tables.extend(page_tables) print(f"Pulled {len(page_tables)} table(s) from this link") else: print(f"No tables found on this page") except requests.exceptions.RequestException as e: print(f"Failed to load link: {e}") except ValueError as e: print(f"Couldn't parse tables from this link: {e}") return all_tables # Example usage (replace these values with your actual URL and heading) if __name__ == "__main__": MAIN_PAGE_URL = "https://your-target-page.com" TARGET_HEADING_TEXT = "Your Specified Heading Here" extracted_tables = extract_tables_under_title(MAIN_PAGE_URL, TARGET_HEADING_TEXT) # Save each table to a separate CSV file for table_num, table in enumerate(extracted_tables, 1): table.to_csv(f"extracted_table_{table_num}.csv", index=False) print(f"Saved table {table_num} to extracted_table_{table_num}.csv")
Step 3: Customize for Your Page
You'll need to tweak a few parts based on your specific webpage:
- Heading tag: If your target title is an
<h3>instead of<h2>, update the tag list in thetarget_headingline. - Link container: If all links are inside a specific div (e.g.,
<div class="resource-links">right after the heading), replace the link-traversal loop with something simpler:# Example: If links are in a dedicated container after the heading link_container = target_heading.find_next_sibling('div', class_='resource-links') if link_container: links = [urljoin(base_url, a['href']) for a in link_container.find_all('a', href=True)] - Dynamic content: If the page loads links via JavaScript (e.g., React/Vue apps), swap
requestswithseleniumto render the page first—just let me know if you need help with that!
内容的提问来源于stack exchange,提问作者user9764822
相关产品推荐
相关产品推荐

