You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python 3提取文本中的国家/城市?pip安装geography失败求助

Hey there! Let's sort out your problem of extracting countries/cities from short texts, and fix that failed library install issue first.

First off, that geography library you tried to install doesn't exist on PyPI at all—so pip can't find any matching versions. No worries though, there are several reliable, well-maintained Python libraries that can do exactly what you need, and they work great with Python 3.

1. Use pycountry for precise country matching

pycountry is a lightweight library packed with official, up-to-date data on countries, regions, and languages. It's perfect for exact matches in your text.

  • Install it (make sure you target Python 3, since you have Python 2 too):
    pip3 install pycountry
    
  • Example code to extract countries:
    import pycountry
    import re
    
    def extract_countries(text):
        # Split text into potential terms (handle parentheses and commas)
        terms = re.split(r'[,()\s]+', text)
        countries = []
        for term in terms:
            term = term.strip()
            if not term:
                continue
            # Try matching official country name
            match = pycountry.countries.get(name=term)
            if not match:
                # Fallback to common names (like "United States" instead of full official name)
                match = pycountry.countries.get(common_name=term)
            if match:
                countries.append(match.name)
        # Remove duplicates and return unique results
        return list(set(countries))
    
    # Test your sample texts
    print(extract_countries("I live in Spain"))  # Output: ['Spain']
    print(extract_countries("United States (New York), United Kingdom (London)"))  # Output: ['United States', 'United Kingdom']
    
2. Use spaCy's NER for flexible location extraction

If you need to pull both countries AND cities from text, spaCy's Named Entity Recognition (NER) is ideal—it automatically spots geographic entities (tagged as GPE, which covers countries, cities, and regions).

  • Install spaCy and the English language model:
    pip3 install spacy
    python3 -m spacy download en_core_web_sm
    
  • Example code:
    import spacy
    
    # Load the pre-trained English NER model
    nlp = spacy.load("en_core_web_sm")
    
    def extract_locations(text):
        doc = nlp(text)
        # Extract all entities labeled as geographic locations
        locations = [ent.text for ent in doc.ents if ent.label_ == "GPE"]
        return locations
    
    # Test cases
    print(extract_locations("I live in Spain"))  # Output: ['Spain']
    print(extract_locations("United States (New York), United Kingdom (London)"))  # Output: ['United States', 'New York', 'United Kingdom', 'London']
    
    If you want to separate countries from cities, you can cross-check the results with pycountry's database to filter out non-country entries.
3. Use geopy for fuzzy location parsing

If your text has vague or incomplete location references, geopy can use geocoding services (like OpenStreetMap's Nominatim) to parse and validate locations. Note: This requires an internet connection.

  • Install geopy:
    pip3 install geopy
    
  • Example code:
    from geopy.geocoders import Nominatim
    
    def extract_country_from_text(text):
        geolocator = Nominatim(user_agent="location_extractor")
        location_results = geolocator.geocode(text, exactly_one=False)
        countries = []
        if location_results:
            for loc in location_results:
                # Pull country name from the address components
                country = loc.raw.get('address', {}).get('country')
                if country and country not in countries:
                    countries.append(country)
        return countries
    
    # Test cases
    print(extract_country_from_text("I live in Spain"))  # Output: ['España'] (map to English with pycountry if needed)
    print(extract_country_from_text("United States (New York), United Kingdom (London)"))  # Output: ['United States of America', 'United Kingdom']
    

One critical reminder: Since you have both Python 2 and 3 installed, always use pip3 for installing packages and python3 for running scripts—this ensures everything uses your Python 3 environment instead of Python 2.

内容的提问来源于stack exchange,提问作者Markus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:36:46