You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改ClinicalTrials.gov爬取代码以获取多值字段?

Fixing Multi-Valued Field Extraction for ClinicalTrials.gov Scraper

Hey there! The issue you're hitting is that your current code overwrites values when multiple tags share the same name (like intervention_type for NCT02170532). Dictionaries can't have duplicate keys, so only the last matching tag's value gets saved. Let's adjust the code to group values by tag name instead.

Here's the revised approach:

Instead of grabbing all tags at once and trying to map them directly to a dictionary, we'll iterate over each field in your subset list, collect all matching values for that field, then join them into a comma-separated string if there are multiple entries.

def clinicalTrialsGov(nctid):
    import requests
    from bs4 import BeautifulSoup
    
    # Fetch and parse the XML
    url = f"https://clinicaltrials.gov/ct2/show/{nctid}?displayxml=true"
    response = requests.get(url)
    data = BeautifulSoup(response.text, "xml")
    
    subset = ['intervention_type', 'study_type', 'allocation', 'intervention_model', 
              'primary_purpose', 'masking', 'enrollment', 'official_title', 'condition', 
              'minimum_age', 'maximum_age', 'gender', 'healthy_volunteers', 'phase', 
              'primary_outcome', 'secondary_outcome', 'number_of_arms']
    
    tag_dict = {}
    for field in subset:
        # Find ALL elements matching the current field
        elements = data.find_all(field)
        # Collect all text values from these elements
        values = [elem.text.strip() for elem in elements]
        # Join with commas if multiple values, else use the single value (or empty string)
        combined_value = ", ".join(values)
        # Create the key with your desired format (ct + capitalized field name)
        key = f"ct{field.capitalize()}"
        tag_dict[key] = combined_value
    
    # Print the results
    for key in tag_dict:
        print(f"{key}: {tag_dict[key]}")

Why this works:

  • For each field in your subset, we explicitly fetch every matching tag using find_all(field), so we don't miss any entries.
  • We collect all text values into a list, then join them with ", " to create the comma-separated string you want for multi-valued fields.
  • Single-valued fields will just have their normal text (no extra commas), so they behave exactly as before.

Testing with NCT02170532:

When you run this code for NCT02170532, the ctIntervention_type entry will now show:
ctIntervention_type: Drug, Drug, Other, Device, Device, Drug
Which matches your desired output perfectly.

内容的提问来源于stack exchange,提问作者jdoe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:07:09