使用BeautifulSoup的find_all_previous抓取指定类别募入決定額数据问题
日本财务省页面特定数据抓取问题解决
需要抓取日本财务省页面中「6.価格競争入札について」和「7.非競争入札について」章节下「(2)募入決定額」对应的数据,但页面元素无清晰层级,现有代码无输出,代码如下:
rows = soup.findAll('span') for cell in r: if "募入決定額" in cell: a=rows[0].find_all_previous("td") for i in a: print(a.get('text'))
现有代码问题分析
- 变量名不匹配:定义了
rows但循环用了未定义的r - 文本判断错误:直接判断字符串是否在元素对象中,应使用
cell.text匹配元素文本内容 - 逻辑偏离目标:用
rows[0]取第一个span的前序td,完全没定位到目标数据区域 - 方法调用错误:
a是td列表,循环内错误调用a.get('text'),应使用i.get_text()或i.text
正确实现代码
import requests from bs4 import BeautifulSoup url = "https://www.mof.go.jp/jgbs/auction/calendar/nyusatsu/resul20211101.htm" response = requests.get(url) response.encoding = "utf-8" soup = BeautifulSoup(response.text, 'html.parser') # 目标章节标题 target_sections = ["6.価格競争入札について", "7.非競争入札について"] for section_title in target_sections: # 定位章节起始元素 section = soup.find('span', text=section_title) if not section: continue print(f"--- {section_title} ---") # 在章节内找到"(2)募入決定額"元素 target_item = section.find_next('span', text=lambda t: t and "(2)募入決定額" in t) if target_item: # 定位募入決定額对应的表格 table = target_item.find_next('table') if table: # 提取表格数据 for row in table.find_all('tr'): cells = row.find_all('td') if cells: print("\t".join([cell.get_text(strip=True) for cell in cells]))
代码逻辑说明
- 请求目标页面并完成解析
- 遍历目标章节标题,定位每个章节的起始位置
- 在章节范围内精准找到「(2)募入決定額」的元素
- 定位该元素后续的表格,遍历提取所有单元格数据
内容的提问来源于stack exchange,提问作者Yuki
相关产品推荐
相关产品推荐

