You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从指定class的div下抓取目标<p>标签内容并实现URL动态格式化?

代码问题分析与优化方案

原代码存在的问题

  • 请求方式错误:页面获取应该用requests.get()而非post(),post通常用于提交数据,get才是请求静态页面的正确方式
  • 循环逻辑错误:遍历每个section-content-item时,每次都调用soup.find('p'),这只会返回整个页面的第一个<p>标签,而非当前div下的内容
  • 未筛选目标区块:没有指定获取第4、5、6个section-content-item(对应索引3、4、5,因为Python列表从0开始计数)
  • 未实现URL动态适配:没有用f-string实现不同元素的URL切换

优化后的代码

import requests
from bs4 import BeautifulSoup as bs

currentPath = "你的当前路径"
print('Current path is:', currentPath)

content_list = []
# 用f-string动态生成目标元素的URL
element = "Antimony"  # 可替换为其他元素名,比如"Gold"
url = f'https://pubchem.ncbi.nlm.nih.gov/element/{element}'       

# 使用get请求获取页面
res = requests.get(url)
res.encoding = 'utf-8'  # 确保编码正确,避免乱码

soup = bs(res.text, 'lxml')

# 获取所有class为section-content-item的div
all_sections = soup.find_all('div', class_="section-content-item")

# 取第4、5、6个区块(对应索引3、4、5)
target_sections = all_sections[3:6]

for section in target_sections:
    # 获取当前区块下的p标签文本
    p_content = section.find('p').get_text(strip=True)
    content_list.append(p_content)
    
print(content_list)

代码说明

  • URL动态适配:通过f'https://pubchem.ncbi.nlm.nih.gov/element/{element}'可以快速替换element变量值,抓取不同元素的页面
  • 目标区块筛选:利用列表切片[3:6]精准获取第4到第6个section-content-item
  • 正确获取区块内容:在循环中调用section.find('p'),确保获取的是当前div下的p标签,而非全局第一个p
  • 编码处理:添加res.encoding = 'utf-8'避免页面文本乱码

内容的提问来源于stack exchange,提问作者user21084142

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 23:55:19