如何高效从PubMed提取全类型临床试验注册ID
Great question! Handling 40+ trial registry IDs doesn't have to mean writing 40 separate regex patterns—here's a scalable, maintainable approach that also maps each ID to its corresponding registry:
First, create a dictionary that stores every registry's ID prefix, regex pattern, and official name. This lets you manage all rules in one place instead of scattering regexes throughout your code:
# Populate this with all 40+ registries using their official ID formats (e.g., from ICTRP) REGISTRY_MAPPINGS = { "NCT": {"pattern": r"NCT\d{8}", "name": "ClinicalTrials.gov"}, "ACTRN": {"pattern": r"ACTRN\d{14}", "name": "Australian New Zealand Clinical Trials Registry"}, "ISRCTN": {"pattern": r"ISRCTN\d{8}", "name": "ISRCTN Registry"}, "ChiCTR": {"pattern": r"ChiCTR\d{8}", "name": "Chinese Clinical Trial Registry"}, "EUCTR": {"pattern": r"EUCTR\d{4}-\d{6}-\d{2}", "name": "EU Clinical Trials Register"}, # Add remaining registries here }
Merge all individual registry patterns into a single regex. This lets you scan text once and capture all valid trial IDs:
import re # Combine patterns into a single regex (named groups help identify which registry matched) combined_pattern = "|".join([f"(?P<{prefix}>{details['pattern']})" for prefix, details in REGISTRY_MAPPINGS.items()])
Write a function that scans relevant fields from your PubMed/MEDLINE data, extracts IDs, and maps each to its registry. This ensures you cover all places trial IDs might appear (abstracts, source identifiers, publication metadata, etc.):
def extract_and_map_trial_ids(pubmed_attributes): # Gather all text fields that could contain trial IDs relevant_fields = [ pubmed_attributes.get("mdAbstract", ""), pubmed_attributes.get("mdSI", ""), pubmed_attributes.get("mdSO", ""), pubmed_attributes.get("mdTitle", "") ] full_text = " ".join(map(str, relevant_fields)) # Find all matching IDs in the text matches = re.finditer(combined_pattern, full_text) # Map each match to its registry and remove duplicates unique_trials = [] seen_ids = set() for match in matches: # Get the matched ID and its corresponding registry prefix for prefix in REGISTRY_MAPPINGS.keys(): trial_id = match.group(prefix) if trial_id and trial_id not in seen_ids: seen_ids.add(trial_id) unique_trials.append({ "trial_id": trial_id, "registry": REGISTRY_MAPPINGS[prefix]["name"] }) break return unique_trials # Example usage with your sample PubMed data sample_attributes = {'mdTitle': 'High-dose versus standard-dose amoxicillin/clavulanate for clinically-diagnosed acute bacterial sinusitis: A randomized clinical trial.', 'mdAbstract': 'BACKGROUND: The recommended treatment for acute bacterial sinusitis in adults, amoxicillin with clavulanate, provides only modest benefit. OBJECTIVE: To see if a higher dose of amoxicillin will lead to more rapid improvement. DESIGN, SETTING, AND PARTICIPANTS: Double-blind randomized trial in which, from November 2014 through February 2017, we enrolled 315 adult outpatients diagnosed with acute sinusitis in accordance with Infectious Disease Society of America guidelines. INTERVENTIONS: Standard-dose (SD) immediate-release (IR) amoxicillin/clavulanate 875 /125 mg (n = 159) vs. high-dose (HD) (n = 156). The original HD formulation, 2000 mg of extended-release (ER) amoxicillin with 125 mg of IR clavulanate twice a day, became unavailable half way through the study. The IRB then approved a revised protocol after patient 180 to provide 1750 mg of IR amoxicillin twice a day in the HD formulation and to compare Time Period 1 (ER) with Time Period 2 (IR). MAIN MEASURE: The primary outcome was the percentage in each group reporting a major improvement-defined as a global assessment of sinusitis symptoms as "a lot better" or "no symptoms"-after 3 days of treatment. KEY RESULTS: Major improvement after 3 days was reported during Period 1 by 38.8% of ER HD versus 37.9% of SD patients (P = 0.91) and during Period 2 by 52.4% of IR HD versus 34.4% of SD patients, an effect size of 18% (95% CI 0.75 to 35%, P = 0.04). No significant differences in efficacy were seen at Day 10. The major side effect, severe diarrhea at Day 3, was reported during Period 1 by 7.4% of HD and 5.7% of SD patients (P = 0.66) and during Period 2 by 15.8% of HD and 4.8% of SD patients (P = 0.048). CONCLUSIONS: Adults with clinically diagnosed acute bacterial sinusitis were more likely to improve rapidly when treated with IR HD than with SD but not when treated with ER HD. They were also more likely to suffer severe diarrhea. Further study is needed to confirm these findings. TRIAL REGISTRATION: ClinicalTrials.gov Identifier: NCT02340000.', 'mdMesh': '', 'mdPMID': '29738561', 'mdPublicationType': ['Journal Article'], 'mdAuthor': ['Matho A', 'Mulqueen M', 'Tanino M', 'Quidort A', 'Cheung J', 'Pollard J', 'Rodriguez J', 'Swamy S', 'Tayler B', 'Garrison G', 'Ata A', 'Sorum P'], 'mdDataPublished': '2018', 'mdPMC': '', 'mdSI': ['ClinicalTrials.gov/NCT02340000'], 'mdAID': ['10.1371/journal.pone.0196734 [doi]', 'PONE-D-17-43190 [pii]'], 'mdDOI': ['10.1371/journal.pone.0196734 [doi]', 'PONE-D-17-43190 [pii]'], 'mdSO': 'PLoS One. 2018 May 8;13(5):e0196734. doi: 10.1371/journal.pone.0196734. eCollection 2018.', 'mdLanguage': ['English']} # Run extraction extracted_trials = extract_and_map_trial_ids(sample_attributes) print(extracted_trials) # Output: [{'trial_id': 'NCT02340000', 'registry': 'ClinicalTrials.gov'}]
- Populate the Full Mapping: Use the WHO ICTRP registry list to get the official regex patterns for all 40+ registries—this ensures you capture every valid ID format.
- Expand Field Coverage: Add other fields from your PubMed data if trial IDs might appear there (e.g.,
mdCommentsif available). - Handle Edge Cases: Some IDs might be wrapped in parentheses or followed by punctuation—you can adjust the regex patterns to include optional cleanup (e.g.,
r"NCT\d{8}[.,)]?"to handle trailing dots/commas/parentheses). - Validate Matches: For extra robustness, you can add a step to cross-check extracted IDs against the registry's official API (if available) to confirm validity.
内容的提问来源于stack exchange,提问作者jdoe

