Python中用BeautifulSoup提取特定Div内容,区分cpRow与cpRow odd
解决方案
1. 分开提取两类标签元素
BeautifulSoup支持通过class_参数精准匹配CSS类名,直接分开获取cpRow和cpRow odd的div元素:
from bs4 import BeautifulSoup from selenium.webdriver import ChromeOptions from selenium import webdriver # 初始化浏览器 options = ChromeOptions() options.add_argument("--headless=new") # 推荐使用新版无头模式 driver = webdriver.Chrome(options) # 目标URL url = "https://cfpub.epa.gov/compliance/criminal_prosecution/index.cfm?action=3&prosecution_summary_id=3420&searchParams=M5%2C%3A%2FXT%2A%5CCYZ%40O%3B%20W%5F%2AYN5%5E%3EK99%2A%29W%3CU%3FV%23DH%5BZ4%247TRPH%3BJQH%229%3FD%3C%26Z%40CY%26%0AM7EFH%21%25%21%3A%23%3DV%40%3A%2A%5F%3AB8%2A%5DR%3BB%25%5E9%5B2D%22I2KE65NEY7M%21%2DU%40%2B8%22J%29Y%23%24LNJ%40DX%24%0A%2F5YJ%3EP%27O%5FK04%5FG%5C%3E%290M4%2E%0A" driver.get(url) soup = BeautifulSoup(driver.page_source, 'html.parser') # 分开提取两类div cp_row_normal = soup.find_all('div', class_='cpRow') cp_row_odd = soup.find_all('div', class_='cpRow odd') driver.quit() # 执行完后关闭浏览器进程
2. 提取指定变量数据
观察页面结构,目标数据分布在这些div中,可通过文本提取和标识判断拆分:
# 提取FISCAL YEAR fiscal_year = None for div in cp_row_normal + cp_row_odd: text = div.get_text(strip=True) if text.startswith('FISCAL YEAR'): fiscal_year = text.split(':')[-1].strip() break # 提取summary和full_text summary = [] full_text = [] is_summary = False is_full_text = False for div in cp_row_normal + cp_row_odd: text = div.get_text(separator=' ', strip=True) if text.startswith('CASE SUMMARY'): is_summary = True is_full_text = False continue elif text.startswith('CASE DETAILS'): is_summary = False is_full_text = True continue if is_summary and text: summary.append(text) elif is_full_text and text: full_text.append(text) # 合并成最终字符串 summary = '\n'.join(summary) full_text = '\n'.join(full_text) # 验证输出 print(f"FISCAL YEAR = {fiscal_year}") print(f"\nsummary = {summary}") print(f"\nfull_text = {full_text}")
关键说明
- 用
class_参数可精准匹配单一CSS类,避免同时获取带odd后缀的元素; - 通过页面标题标识(CASE SUMMARY/CASE DETAILS)区分数据归属,确保内容准确拆分到对应变量;
- 加入
driver.quit()避免浏览器进程残留。
内容的提问来源于stack exchange,提问作者emiley mille
相关产品推荐
相关产品推荐

