如何从bs4.element.ResultSet中提取Google Scholar XML记录里jats:abstract标签下的纯文本内容
Got it, let's break down how to get that clean, tag-free abstract text you need.
First, remember that soup.find_all("jats:abstract") returns a bs4.element.ResultSet—a list-like object containing all matching <jats:abstract> elements. You can't just call .get_text() directly on the whole ResultSet; you need to loop through each element in it first.
Step-by-Step Solution
- Iterate over the ResultSet: Loop through each abstract element in your
xml_abstractvariable. - Extract pure text: Use BeautifulSoup's
.get_text()method on each element. This automatically strips all HTML/XML tags, and you can tweak parameters to clean up whitespace.
Here's the code:
# Assuming xml_abstract is your ResultSet from soup.find_all("jats:abstract") clean_abstracts = [] for abstract_element in xml_abstract: # Extract text, strip extra whitespace, and join with single spaces plain_text = abstract_element.get_text(strip=True, separator=' ') clean_abstracts.append(plain_text) # If you only expect one abstract, you can grab it directly: # plain_text = xml_abstract[0].get_text(strip=True, separator=' ')
What the Parameters Do
strip=True: Removes leading/trailing whitespace from the entire text block.separator=' ': Ensures text from different child tags (like<jats:title>and<jats:p>) is separated by a single space, so you don't get messy concatenation like "SummaryEstablished primary prevention...".
Example Output
Running this code on your sample XML will give you this clean plain text:
Summary Established primary prevention strategies of cardiovascular diseases are based on understanding of risk factors, but whether the same risk factors are associated with atrial fibrillation (AF) remains unclear. We conducted a systematic review and field synopsis of the associations of 23 cardiovascular risk factors and incident AF, which included 84 reports based on 28 consented and four electronic health record cohorts of 20,420,175 participants and 576,602 AF events. We identified 3-19 reports per risk factor and heterogeneity in AF definition, quality of reporting, and adjustment. We extracted relative risks (RR) and 95 % confidence intervals [CI] and visualised the number of reports with inverse (RR [CI]<1.00), or direct (RR [CI]>1.00) associations. For hypertension (13/17 reports) and obesity (19/19 reports), there were direct associations with incident AF, as there are for coronary heart disease (CHD). There were inverse associations for non-White ethnicity (5/5 reports, with RR from 0.35 to 0.84 [0.82–0.85]), total cholesterol (4/13 reports from 0.76 [0.59–0.98] to 0.94 [0.90–0.97]; 8/13 reports with non-significant inverse associations), and diastolic blood pressure (2/11 reports from 0.87 [0.78–0.96] to 0.92 [0.85–0.99]; 5/11 reports with non-significant inverse associations), and direct associations for taller height (7/10 reports from 1.03 [1.02–1.05] to 1.92 [1.38–2.67]), which are in the opposite direction of known associations with CHD. A systematic evaluation of the available evidence suggests similarities as well as important differences in the risk factors for incidence of AF as compared with other cardiovascular diseases, which has implications for the primary prevention strategies for atrial fibrillation.
内容的提问来源于stack exchange,提问作者Giorgia Mancini

