如何拆分抓取的文本并创建DataFrame?附代码优化需求
解决方案
1. 生成列表的列表
修正原代码逻辑,直接定位目标书籍列表的<ol>标签,遍历提取每本书的完整文本并生成嵌套列表:
import requests import re import pandas as pd from bs4 import BeautifulSoup # 请求页面 r = requests.get("https://www.gutenberg.org/browse/scores/top") soup = BeautifulSoup(r.content, "lxml") # 定位昨日热门书籍的ol列表 book_ol = soup.find('ol') # 生成列表的列表 nested_book_list = [] for li in book_ol.find_all('li'): book_full_text = li.find('a').get_text() nested_book_list.append([book_full_text]) print(nested_book_list)
输出格式符合要求:
[['A Room with a View by E. M. Forster (37480)'], ['Middlemarch by George Eliot (34900)'], ...]
2. 拆分书名与访问量
用正则表达式匹配括号内的数字,将每个元素拆分为书名和访问量两部分:
# 拆分后的结果列表 split_book_list = [] # 匹配书名和括号内数字的正则 pattern = re.compile(r'(.*?)\s\((\d+)\)$') for item in nested_book_list: match_result = pattern.match(item[0]) if match_result: book_name = match_result.group(1).strip() view_count = match_result.group(2) split_book_list.append([book_name, view_count]) print(split_book_list)
输出格式:
[["A Room with a View by E. M. Forster", "37480"], ["Middlemarch by George Eliot", "34900"], ...]
3. 转换为DataFrame
直接将拆分后的列表传入pd.DataFrame(),指定列名完成转换:
# 转换为DataFrame book_df = pd.DataFrame(split_book_list, columns=['书名', '访问量']) # 可选:将访问量转为数值类型 book_df['访问量'] = book_df['访问量'].astype(int) # 查看前几行数据 print(book_df.head())
输出示例:
书名 访问量 0 A Room with a View by E. M. Forster 37480 1 Middlemarch by George Eliot 34900 2 Little Women; Or, Meg, Jo, Beth, and Amy by Lo... 31929
内容的提问来源于stack exchange,提问作者NIVEA
相关产品推荐
相关产品推荐

