如何抓取网页生成以子标题为列名、li内容为行的DataFrame
问题原因
原代码独立提取标题和ul列表,没有建立标题与对应下方内容的映射关系,导致列名和内容错位,且未导入pandas库、提前创建固定长度空DataFrame的方式灵活性不足。
修正后完整代码
from bs4 import BeautifulSoup import requests import pandas as pd # 增加请求头避免反爬拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } url = "https://www.bankersadda.com/17th-september-2021-daily-gk-update/" page = requests.get(url, headers=headers) page.encoding = page.apparent_encoding if page.status_code != 200: raise Exception("页面请求失败,请检查链接或网络") soup = BeautifulSoup(page.text, 'lxml') article = soup.find(class_ = "entry-content") # 用字典存储标题与对应内容,自动建立映射 data_dict = {} current_heading = None # 遍历正文区域所有子节点,匹配标题和对应内容 for child in article.children: if child.name == 'p': strong_tag = child.find('strong') # 匹配带News的标题 if strong_tag and 'News' in strong_tag.text.strip(): current_heading = strong_tag.text.strip() data_dict[current_heading] = [] # 匹配当前标题下方对应的ul内容 elif child.name == 'ul' and current_heading is not None: for li in child.find_all('li'): data_dict[current_heading].append(li.text.strip()) # 匹配完当前标题的内容后重置标记,等待下一个标题 current_heading = None # 转换为DataFrame,自动补全不同列的空值为NaN my_df = pd.DataFrame.from_dict(data_dict, orient='index').T print(my_df)
逻辑说明
- 遍历正文所有子节点,遇到带News的标题就记录为当前激活的标题,后续遇到的第一个ul自动归到该标题下
- 字典结构天然实现标题和内容的绑定,不会出现错位问题
- 最后转换为DataFrame时自动处理不同列内容长度不一致的问题,适配页面不规则的内容结构
内容的提问来源于stack exchange,提问作者MRD
相关产品推荐
相关产品推荐

