You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从bs4.element.ResultSet中提取Google Scholar XML记录里jats:abstract标签下的纯文本内容

Extracting Plain Text from jats:abstract Tags in BeautifulSoup

Got it, let's break down how to get that clean, tag-free abstract text you need.

First, remember that soup.find_all("jats:abstract") returns a bs4.element.ResultSet—a list-like object containing all matching <jats:abstract> elements. You can't just call .get_text() directly on the whole ResultSet; you need to loop through each element in it first.

Step-by-Step Solution

  1. Iterate over the ResultSet: Loop through each abstract element in your xml_abstract variable.
  2. Extract pure text: Use BeautifulSoup's .get_text() method on each element. This automatically strips all HTML/XML tags, and you can tweak parameters to clean up whitespace.

Here's the code:

# Assuming xml_abstract is your ResultSet from soup.find_all("jats:abstract")
clean_abstracts = []
for abstract_element in xml_abstract:
    # Extract text, strip extra whitespace, and join with single spaces
    plain_text = abstract_element.get_text(strip=True, separator=' ')
    clean_abstracts.append(plain_text)

# If you only expect one abstract, you can grab it directly:
# plain_text = xml_abstract[0].get_text(strip=True, separator=' ')

What the Parameters Do

  • strip=True: Removes leading/trailing whitespace from the entire text block.
  • separator=' ': Ensures text from different child tags (like <jats:title> and <jats:p>) is separated by a single space, so you don't get messy concatenation like "SummaryEstablished primary prevention...".

Example Output

Running this code on your sample XML will give you this clean plain text:

Summary Established primary prevention strategies of cardiovascular diseases are based on understanding of risk factors, but whether the same risk factors are associated with atrial fibrillation (AF) remains unclear. We conducted a systematic review and field synopsis of the associations of 23 cardiovascular risk factors and incident AF, which included 84 reports based on 28 consented and four electronic health record cohorts of 20,420,175 participants and 576,602 AF events. We identified 3-19 reports per risk factor and heterogeneity in AF definition, quality of reporting, and adjustment. We extracted relative risks (RR) and 95 % confidence intervals [CI] and visualised the number of reports with inverse (RR [CI]<1.00), or direct (RR [CI]>1.00) associations. For hypertension (13/17 reports) and obesity (19/19 reports), there were direct associations with incident AF, as there are for coronary heart disease (CHD). There were inverse associations for non-White ethnicity (5/5 reports, with RR from 0.35 to 0.84 [0.82–0.85]), total cholesterol (4/13 reports from 0.76 [0.59–0.98] to 0.94 [0.90–0.97]; 8/13 reports with non-significant inverse associations), and diastolic blood pressure (2/11 reports from 0.87 [0.78–0.96] to 0.92 [0.85–0.99]; 5/11 reports with non-significant inverse associations), and direct associations for taller height (7/10 reports from 1.03 [1.02–1.05] to 1.92 [1.38–2.67]), which are in the opposite direction of known associations with CHD. A systematic evaluation of the available evidence suggests similarities as well as important differences in the risk factors for incidence of AF as compared with other cardiovascular diseases, which has implications for the primary prevention strategies for atrial fibrillation.

内容的提问来源于stack exchange,提问作者Giorgia Mancini

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.01 01:07:30