如何用Python和Beautiful Soup提取本地HTML表格的标题与URL?
没问题,我帮你搞定这个需求!用Python+BeautifulSoup其实不难,我给你写个清晰的实现方案,你可以根据自己的HTML结构调整细节:
第一步:先安装必要的库
如果你还没装BeautifulSoup,先在终端里跑这个命令:
pip install beautifulsoup4
第二步:核心代码实现
下面分两种常见的HTML结构场景给你写代码,你对应自己的文件选就行:
场景1:表格标题在<caption>标签里(比如表格自带标题)
from bs4 import BeautifulSoup # 读取本地的HTML文件 with open('你的文件名.html', 'r', encoding='utf-8') as file: soup = BeautifulSoup(file, 'html.parser') # 初始化Markdown表格的表头 markdown_result = "| 标题 | URL |\n|------|-----|\n" # 遍历页面里的每一个表格 for table in soup.find_all('table'): # 提取表格标题,没有标题的话显示"无标题" table_title = table.caption.get_text(strip=True) if table.caption else "无标题" # 提取表格里的所有URL,用集合去重避免重复链接 url_set = set() for link in table.find_all('a', href=True): url = link['href'] url_set.add(url) # 把标题和对应的URL逐个加到Markdown表格里 for url in url_set: markdown_result += f"| {table_title} | {url} |\n" # 打印结果,或者保存到Markdown文件里 print(markdown_result) # 保存到本地文件 with open('整理结果.md', 'w', encoding='utf-8') as md_file: md_file.write(markdown_result)
场景2:表格标题在表格上方的标题标签里(比如<h2>/<h3>)
如果你的标题不是表格自带的,而是表格前面的<h2>(比如示例里的"Windows"是一个<h2>),那把提取标题的部分改一下就行:
from bs4 import BeautifulSoup with open('你的文件名.html', 'r', encoding='utf-8') as file: soup = BeautifulSoup(file, 'html.parser') markdown_result = "| 标题 | URL |\n|------|-----|\n" for table in soup.find_all('table'): # 找表格前面最近的<h2>标签作为标题,你可以改成h3/h4对应自己的结构 title_tag = table.find_previous('h2') table_title = title_tag.get_text(strip=True) if title_tag else "无标题" # 下面提取URL的部分和上面一样 url_set = set() for link in table.find_all('a', href=True): url_set.add(link['href']) for url in url_set: markdown_result += f"| {table_title} | {url} |\n" print(markdown_result) # 保存到文件 with open('整理结果.md', 'w', encoding='utf-8') as md_file: md_file.write(markdown_result)
一些适配小技巧
- 如果你的URL只在表格的某一列里(比如只有第二列有链接),可以把
table.find_all('a')改成table.find_all('td', class_='url-column')(替换成你实际的类名)再找链接,更精准。 - 如果不需要去重,把
set()改成普通列表就行。 - 遇到中文乱码的话,确保打开文件时指定了
encoding='utf-8',这个很重要!
内容的提问来源于stack exchange,提问作者user9796664
相关产品推荐
相关产品推荐

