Python爬取BIS演讲列表时无法正确提取详情页正文问题求助
修复方案
代码跑不通是三个明确问题,对应修改后就能拿到包含date、author、title、text四列的目标dataframe:
- 链接选择器写法错误:你遍历的
<tr>标签本身没有href属性,实际详情页地址存放在每行标题对应的<a>标签上,直接读取tr的属性拿不到有效跳转地址 - 正文选择范围过大:直接匹配全页面所有
<p>标签会把导航栏、页脚、相关推荐等无关内容全部抓入,BIS的演讲正文统一存放在class="cms-content"的容器内,仅抓取该容器下的段落即可拿到纯正文 - 日期字段提取逻辑缺失:之前写的
card.select('.item_date')仅获取了标签对象,没有提取标签内的文本,最终存入数据的不是实际日期字符串
修正后的完整可运行代码:
from bs4 import BeautifulSoup import requests import pandas as pd payload = 'from=&till=&objid=cbspeeches&page=&paging_length=10&sort_list=date_desc&theme=cbspeeches&ml=false&mlurl=&emptylisttext=' url = 'https://www.bis.org/doclist/cbspeeches.htm' headers = { "content-type": "application/x-www-form-urlencoded", "X-Requested-With": "XMLHttpRequest", "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36" } req = requests.post(url, headers=headers, data=payload) soup = BeautifulSoup(req.content, "lxml") data = [] for card in soup.select('.documentList tbody tr'): # 提取日期文本 date = card.select_one('.item_date').get_text(strip=True) title_tag = card.select_one('.title a') title = title_tag.get_text(strip=True) author = card.select_one('.authorlnk.dashed').get_text(strip=True) # 拼接详情页完整地址 detail_url = f"https://www.bis.org{title_tag['href']}" # 请求并解析详情页 detail_resp = requests.get(detail_url, headers=headers) detail_soup = BeautifulSoup(detail_resp.content, "lxml") # 提取正文纯文本,段落用换行分隔 text = '\n'.join([p.get_text(strip=True) for p in detail_soup.select('.cms-content p') if p.get_text(strip=True)]) data.append({ 'date': date, 'author': author, 'title': title, 'text': text }) # 转为目标dataframe df = pd.DataFrame(data) print(df)
补充说明:
- 新增了常规浏览器标识的
User-Agent请求头,避免触发网站反爬规则返回403错误 - 正文提取加了空段落过滤,不会存入无意义的空白行
- 所有文本提取都加了空白裁剪,自动去除首尾多余的空格、换行符
- 把列表字段提取和详情页请求合并到同一次循环,不会覆盖之前已提取的字段,不需要重复解析页面
内容的提问来源于stack exchange,提问作者Rollo99
相关产品推荐
相关产品推荐

