You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中用BeautifulSoup提取特定Div内容,区分cpRow与cpRow odd

解决方案

1. 分开提取两类标签元素

BeautifulSoup支持通过class_参数精准匹配CSS类名,直接分开获取cpRow和cpRow odd的div元素:

from bs4 import BeautifulSoup
from selenium.webdriver import ChromeOptions
from selenium import webdriver

# 初始化浏览器
options = ChromeOptions()
options.add_argument("--headless=new")  # 推荐使用新版无头模式
driver = webdriver.Chrome(options)

# 目标URL
url = "https://cfpub.epa.gov/compliance/criminal_prosecution/index.cfm?action=3&prosecution_summary_id=3420&searchParams=M5%2C%3A%2FXT%2A%5CCYZ%40O%3B%20W%5F%2AYN5%5E%3EK99%2A%29W%3CU%3FV%23DH%5BZ4%247TRPH%3BJQH%229%3FD%3C%26Z%40CY%26%0AM7EFH%21%25%21%3A%23%3DV%40%3A%2A%5F%3AB8%2A%5DR%3BB%25%5E9%5B2D%22I2KE65NEY7M%21%2DU%40%2B8%22J%29Y%23%24LNJ%40DX%24%0A%2F5YJ%3EP%27O%5FK04%5FG%5C%3E%290M4%2E%0A"
driver.get(url)
soup = BeautifulSoup(driver.page_source, 'html.parser')

# 分开提取两类div
cp_row_normal = soup.find_all('div', class_='cpRow')
cp_row_odd = soup.find_all('div', class_='cpRow odd')

driver.quit()  # 执行完后关闭浏览器进程

2. 提取指定变量数据

观察页面结构,目标数据分布在这些div中,可通过文本提取和标识判断拆分:

# 提取FISCAL YEAR
fiscal_year = None
for div in cp_row_normal + cp_row_odd:
    text = div.get_text(strip=True)
    if text.startswith('FISCAL YEAR'):
        fiscal_year = text.split(':')[-1].strip()
        break

# 提取summary和full_text
summary = []
full_text = []
is_summary = False
is_full_text = False

for div in cp_row_normal + cp_row_odd:
    text = div.get_text(separator=' ', strip=True)
    if text.startswith('CASE SUMMARY'):
        is_summary = True
        is_full_text = False
        continue
    elif text.startswith('CASE DETAILS'):
        is_summary = False
        is_full_text = True
        continue
    
    if is_summary and text:
        summary.append(text)
    elif is_full_text and text:
        full_text.append(text)

# 合并成最终字符串
summary = '\n'.join(summary)
full_text = '\n'.join(full_text)

# 验证输出
print(f"FISCAL YEAR = {fiscal_year}")
print(f"\nsummary = {summary}")
print(f"\nfull_text = {full_text}")

关键说明

  • 用class_参数可精准匹配单一CSS类,避免同时获取带odd后缀的元素;
  • 通过页面标题标识(CASE SUMMARY/CASE DETAILS)区分数据归属,确保内容准确拆分到对应变量;
  • 加入driver.quit()避免浏览器进程残留。

内容的提问来源于stack exchange,提问作者emiley mille

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 02:17:15