Excel抓取网页带超链接数据咨询:如何获取含超链接的原格式表格
Absolutely! Grabbing tables with hyperlinks and their original structure intact is totally doable—you just need to tweak how you extract and structure the data instead of only pulling plain text. Let me walk you through a couple of reliable approaches using common Python tools, since they’re the go-to for web scraping tasks:
1. For Static Webpages (BeautifulSoup + Pandas)
If the table is rendered directly in the page’s HTML (no JavaScript needed to load it), this combo works perfectly. We’ll parse the page, check each table cell for hyperlinks, and build structured content that preserves those links.
Here’s a working example:
from bs4 import BeautifulSoup import pandas as pd import requests from urllib.parse import urljoin # For fixing relative links # Replace with your target webpage URL target_url = "https://example.com/your-table-page" response = requests.get(target_url) soup = BeautifulSoup(response.text, "html.parser") # Locate your target table (adjust the selector if needed, e.g., by class or ID) target_table = soup.find("table", class_="your-table-class") # Process each row and cell processed_rows = [] for row in target_table.find_all("tr"): current_row = [] # Iterate through both header cells (th) and data cells (td) for cell in row.find_all(["th", "td"]): link = cell.find("a") if link: # Extract link text and URL (fix relative links to absolute) link_text = link.get_text(strip=True) absolute_url = urljoin(target_url, link["href"]) # Format as a Markdown link (or use HTML if you prefer) formatted_link = f"[{link_text}]({absolute_url})" current_row.append(formatted_link) else: # No link? Just add the cell's plain text current_row.append(cell.get_text(strip=True)) processed_rows.append(current_row) # Convert to a DataFrame and save/output table_df = pd.DataFrame(processed_rows[1:], columns=processed_rows[0]) # Assume first row is headers # Option 1: Save as a Markdown table (preserves links in readable format) with open("scraped_table.md", "w") as md_file: md_file.write(table_df.to_markdown(index=False)) # Option 2: Save as HTML to retain original table styling (if any) table_df.to_html("scraped_table.html", escape=False, index=False)
Key Notes for This Method:
- Use
urljointo convert relative links (like/page/1) to full absolute URLs—this ensures the links work when you open the saved file. - If you want to keep the table’s original CSS styling, saving as HTML is better than Markdown.
2. For Dynamic Webpages (Selenium + BeautifulSoup)
If the table loads after JavaScript runs (e.g., it’s fetched via an API when the page loads), BeautifulSoup alone can’t access it. Use Selenium to simulate a browser loading the page fully, then parse the rendered HTML.
Example code:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup import pandas as pd from urllib.parse import urljoin # Set up headless browser (so no window pops up) chrome_options = Options() chrome_options.add_argument("--headless=new") driver = webdriver.Chrome(options=chrome_options) target_url = "https://example.com/dynamic-table-page" driver.get(target_url) # Wait a few seconds if needed (for JS to finish loading the table) # driver.implicitly_wait(5) # Parse the fully rendered page source soup = BeautifulSoup(driver.page_source, "html.parser") driver.quit() # Same processing as the static method from here target_table = soup.find("table") processed_rows = [] for row in target_table.find_all("tr"): current_row = [] for cell in row.find_all(["th", "td"]): link = cell.find("a") if link: link_text = link.get_text(strip=True) absolute_url = urljoin(target_url, link["href"]) formatted_link = f"[{link_text}]({absolute_url})" current_row.append(formatted_link) else: current_row.append(cell.get_text(strip=True)) processed_rows.append(current_row) table_df = pd.DataFrame(processed_rows[1:], columns=processed_rows[0]) table_df.to_markdown("dynamic_scraped_table.md", index=False)
Quick Tips:
- Always check the website’s
robots.txtfile (e.g.,https://example.com/robots.txt) to make sure scraping is allowed. - If the table has merged cells (colspan/rowspan), Pandas’
to_markdownmight not handle them perfectly—saving as HTML will preserve that structure better.
内容的提问来源于stack exchange,提问作者FaRz1 Ezpz

