You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何抓取网页生成以子标题为列名、li内容为行的DataFrame

问题原因

原代码独立提取标题和ul列表,没有建立标题与对应下方内容的映射关系,导致列名和内容错位,且未导入pandas库、提前创建固定长度空DataFrame的方式灵活性不足。

修正后完整代码

from bs4 import BeautifulSoup
import requests
import pandas as pd

# 增加请求头避免反爬拦截
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}
url = "https://www.bankersadda.com/17th-september-2021-daily-gk-update/"
page = requests.get(url, headers=headers)
page.encoding = page.apparent_encoding
if page.status_code != 200:
    raise Exception("页面请求失败,请检查链接或网络")

soup = BeautifulSoup(page.text, 'lxml')
article = soup.find(class_ = "entry-content")

# 用字典存储标题与对应内容,自动建立映射
data_dict = {}
current_heading = None

# 遍历正文区域所有子节点,匹配标题和对应内容
for child in article.children:
    if child.name == 'p':
        strong_tag = child.find('strong')
        # 匹配带News的标题
        if strong_tag and 'News' in strong_tag.text.strip():
            current_heading = strong_tag.text.strip()
            data_dict[current_heading] = []
    # 匹配当前标题下方对应的ul内容
    elif child.name == 'ul' and current_heading is not None:
        for li in child.find_all('li'):
            data_dict[current_heading].append(li.text.strip())
        # 匹配完当前标题的内容后重置标记,等待下一个标题
        current_heading = None

# 转换为DataFrame,自动补全不同列的空值为NaN
my_df = pd.DataFrame.from_dict(data_dict, orient='index').T
print(my_df)

逻辑说明

  • 遍历正文所有子节点,遇到带News的标题就记录为当前激活的标题,后续遇到的第一个ul自动归到该标题下
  • 字典结构天然实现标题和内容的绑定,不会出现错位问题
  • 最后转换为DataFrame时自动处理不同列内容长度不一致的问题,适配页面不规则的内容结构

内容的提问来源于stack exchange,提问作者MRD

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 07:39:04