You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Excel抓取网页带超链接数据咨询:如何获取含超链接的原格式表格

Absolutely! Grabbing tables with hyperlinks and their original structure intact is totally doable—you just need to tweak how you extract and structure the data instead of only pulling plain text. Let me walk you through a couple of reliable approaches using common Python tools, since they’re the go-to for web scraping tasks:

1. For Static Webpages (BeautifulSoup + Pandas)

If the table is rendered directly in the page’s HTML (no JavaScript needed to load it), this combo works perfectly. We’ll parse the page, check each table cell for hyperlinks, and build structured content that preserves those links.

Here’s a working example:

from bs4 import BeautifulSoup
import pandas as pd
import requests
from urllib.parse import urljoin  # For fixing relative links

# Replace with your target webpage URL
target_url = "https://example.com/your-table-page"
response = requests.get(target_url)
soup = BeautifulSoup(response.text, "html.parser")

# Locate your target table (adjust the selector if needed, e.g., by class or ID)
target_table = soup.find("table", class_="your-table-class")

# Process each row and cell
processed_rows = []
for row in target_table.find_all("tr"):
    current_row = []
    # Iterate through both header cells (th) and data cells (td)
    for cell in row.find_all(["th", "td"]):
        link = cell.find("a")
        if link:
            # Extract link text and URL (fix relative links to absolute)
            link_text = link.get_text(strip=True)
            absolute_url = urljoin(target_url, link["href"])
            # Format as a Markdown link (or use HTML if you prefer)
            formatted_link = f"[{link_text}]({absolute_url})"
            current_row.append(formatted_link)
        else:
            # No link? Just add the cell's plain text
            current_row.append(cell.get_text(strip=True))
    processed_rows.append(current_row)

# Convert to a DataFrame and save/output
table_df = pd.DataFrame(processed_rows[1:], columns=processed_rows[0])  # Assume first row is headers

# Option 1: Save as a Markdown table (preserves links in readable format)
with open("scraped_table.md", "w") as md_file:
    md_file.write(table_df.to_markdown(index=False))

# Option 2: Save as HTML to retain original table styling (if any)
table_df.to_html("scraped_table.html", escape=False, index=False)

Key Notes for This Method:

  • Use urljoin to convert relative links (like /page/1) to full absolute URLs—this ensures the links work when you open the saved file.
  • If you want to keep the table’s original CSS styling, saving as HTML is better than Markdown.

2. For Dynamic Webpages (Selenium + BeautifulSoup)

If the table loads after JavaScript runs (e.g., it’s fetched via an API when the page loads), BeautifulSoup alone can’t access it. Use Selenium to simulate a browser loading the page fully, then parse the rendered HTML.

Example code:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup
import pandas as pd
from urllib.parse import urljoin

# Set up headless browser (so no window pops up)
chrome_options = Options()
chrome_options.add_argument("--headless=new")
driver = webdriver.Chrome(options=chrome_options)

target_url = "https://example.com/dynamic-table-page"
driver.get(target_url)

# Wait a few seconds if needed (for JS to finish loading the table)
# driver.implicitly_wait(5)

# Parse the fully rendered page source
soup = BeautifulSoup(driver.page_source, "html.parser")
driver.quit()

# Same processing as the static method from here
target_table = soup.find("table")
processed_rows = []
for row in target_table.find_all("tr"):
    current_row = []
    for cell in row.find_all(["th", "td"]):
        link = cell.find("a")
        if link:
            link_text = link.get_text(strip=True)
            absolute_url = urljoin(target_url, link["href"])
            formatted_link = f"[{link_text}]({absolute_url})"
            current_row.append(formatted_link)
        else:
            current_row.append(cell.get_text(strip=True))
    processed_rows.append(current_row)

table_df = pd.DataFrame(processed_rows[1:], columns=processed_rows[0])
table_df.to_markdown("dynamic_scraped_table.md", index=False)

Quick Tips:

  • Always check the website’s robots.txt file (e.g., https://example.com/robots.txt) to make sure scraping is allowed.
  • If the table has merged cells (colspan/rowspan), Pandas’ to_markdown might not handle them perfectly—saving as HTML will preserve that structure better.

内容的提问来源于stack exchange,提问作者FaRz1 Ezpz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 06:35:48