如何遍历HTML文件并将特定数据解析至DataFrame?
问题:将docx转换的混乱HTML解析为DataFrame表格
我试过BeautifulSoup、XML解析器等多种方法,想找个更简便的方式遍历HTML文件,把信息解析成DataFrame。这个HTML是docx转来的,格式很乱,只需要把<b>数字</b>这类加粗数字后面的每段文本拆成单独行。
示例HTML代码
<h2 class="chapter-header-western">CHAPTER 1</h2> <p class="western" style="line-height: 100%; margin-bottom: 0.08in"><b>1</b>text <p class="western" style="line-height: 100%; margin-bottom: 0.08in"><b>Header</b></p> <p align="left" style="line-height: 100%; margin-bottom: 0.08in"><b>2</b>text </p> <p align="left" style="line-height: 120%; margin-left: 0.3in; text-indent: -0.3in; margin-bottom: 0.08in"> <b>3 </b>text </p> <p class="western" style="line-height: 100%; margin-bottom: 0.08in">text <b>4 </b>text <b>5 </b>text<b>6 </b>text </p>
目标DataFrame格式
| Chapter | Number | Text |
|---|---|---|
| 1 | 1 | text |
| 1 | 2 | text |
| 1 | 3 | text |
我试过BeautifulSoup的find_all方法,但它只返回标签内的字符串,我需要的是标签之后的文本。或许可以把<b>数字</b>作为分隔标记来处理?
内容的提问来源于stack exchange,提问作者user16950345
相关产品推荐
相关产品推荐

