You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python和Beautiful Soup提取本地HTML表格的标题与URL?

没问题,我帮你搞定这个需求!用Python+BeautifulSoup其实不难,我给你写个清晰的实现方案,你可以根据自己的HTML结构调整细节:

第一步:先安装必要的库

如果你还没装BeautifulSoup,先在终端里跑这个命令:

pip install beautifulsoup4

第二步:核心代码实现

下面分两种常见的HTML结构场景给你写代码,你对应自己的文件选就行:

场景1:表格标题在<caption>标签里(比如表格自带标题)

from bs4 import BeautifulSoup

# 读取本地的HTML文件
with open('你的文件名.html', 'r', encoding='utf-8') as file:
    soup = BeautifulSoup(file, 'html.parser')

# 初始化Markdown表格的表头
markdown_result = "| 标题 | URL |\n|------|-----|\n"

# 遍历页面里的每一个表格
for table in soup.find_all('table'):
    # 提取表格标题,没有标题的话显示"无标题"
    table_title = table.caption.get_text(strip=True) if table.caption else "无标题"
    
    # 提取表格里的所有URL,用集合去重避免重复链接
    url_set = set()
    for link in table.find_all('a', href=True):
        url = link['href']
        url_set.add(url)
    
    # 把标题和对应的URL逐个加到Markdown表格里
    for url in url_set:
        markdown_result += f"| {table_title} | {url} |\n"

# 打印结果,或者保存到Markdown文件里
print(markdown_result)
# 保存到本地文件
with open('整理结果.md', 'w', encoding='utf-8') as md_file:
    md_file.write(markdown_result)

场景2:表格标题在表格上方的标题标签里(比如<h2>/<h3>)

如果你的标题不是表格自带的,而是表格前面的<h2>(比如示例里的"Windows"是一个<h2>),那把提取标题的部分改一下就行:

from bs4 import BeautifulSoup

with open('你的文件名.html', 'r', encoding='utf-8') as file:
    soup = BeautifulSoup(file, 'html.parser')

markdown_result = "| 标题 | URL |\n|------|-----|\n"

for table in soup.find_all('table'):
    # 找表格前面最近的<h2>标签作为标题,你可以改成h3/h4对应自己的结构
    title_tag = table.find_previous('h2')
    table_title = title_tag.get_text(strip=True) if title_tag else "无标题"
    
    # 下面提取URL的部分和上面一样
    url_set = set()
    for link in table.find_all('a', href=True):
        url_set.add(link['href'])
    
    for url in url_set:
        markdown_result += f"| {table_title} | {url} |\n"

print(markdown_result)
# 保存到文件
with open('整理结果.md', 'w', encoding='utf-8') as md_file:
    md_file.write(markdown_result)

一些适配小技巧

  • 如果你的URL只在表格的某一列里(比如只有第二列有链接),可以把table.find_all('a')改成table.find_all('td', class_='url-column')(替换成你实际的类名)再找链接,更精准。
  • 如果不需要去重,把set()改成普通列表就行。
  • 遇到中文乱码的话,确保打开文件时指定了encoding='utf-8',这个很重要!

内容的提问来源于stack exchange,提问作者user9796664

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 10:07:24