如何根据交替排列的标题与文本内容创建指定格式的DataFrame
实现方案
使用BeautifulSoup解析HTML提取所有段落文本,再按标题、文本交替出现的规则分组后构建DataFrame即可,具体操作如下:
首先安装所需依赖(已安装可跳过):pip install pandas beautifulsoup4
完整实现代码:
import pandas as pd from bs4 import BeautifulSoup # 此处可替换为从本地HTML文件读取的内容 html_content = """ <p>Heading 1</p> <p>Some text here</p> <p>Heading 2</p> <p>Some text here</p> <p>Heading 3</p> <p>Some text here</p> """ # 解析提取所有p标签的文本 soup = BeautifulSoup(html_content, "html.parser") all_p_text = [p.get_text(strip=True) for p in soup.find_all("p")] # 按规则拆分:索引偶数为标题,索引奇数为对应文本 df = pd.DataFrame({ "Heading": all_p_text[::2], "text": all_p_text[1::2] }) # 打印验证结果 print(df)
如果需要从本地HTML文件读取内容,将html_content赋值部分替换为以下代码即可:
with open("你的本地HTML文件路径.html", "r", encoding="utf-8") as f: html_content = f.read()
运行后输出的df结构如下:
| Heading | text |
|---|---|
| Heading 1 | Some text here |
| Heading 2 | Some text here |
| Heading 3 | Some text here |
内容的提问来源于stack exchange,提问作者arathi
相关产品推荐
相关产品推荐

