如何用Python 3提取文本中的国家/城市?pip安装geography失败求助
Hey there! Let's sort out your problem of extracting countries/cities from short texts, and fix that failed library install issue first.
First off, that geography library you tried to install doesn't exist on PyPI at all—so pip can't find any matching versions. No worries though, there are several reliable, well-maintained Python libraries that can do exactly what you need, and they work great with Python 3.
pycountry is a lightweight library packed with official, up-to-date data on countries, regions, and languages. It's perfect for exact matches in your text.
- Install it (make sure you target Python 3, since you have Python 2 too):
pip3 install pycountry - Example code to extract countries:
import pycountry import re def extract_countries(text): # Split text into potential terms (handle parentheses and commas) terms = re.split(r'[,()\s]+', text) countries = [] for term in terms: term = term.strip() if not term: continue # Try matching official country name match = pycountry.countries.get(name=term) if not match: # Fallback to common names (like "United States" instead of full official name) match = pycountry.countries.get(common_name=term) if match: countries.append(match.name) # Remove duplicates and return unique results return list(set(countries)) # Test your sample texts print(extract_countries("I live in Spain")) # Output: ['Spain'] print(extract_countries("United States (New York), United Kingdom (London)")) # Output: ['United States', 'United Kingdom']
If you need to pull both countries AND cities from text, spaCy's Named Entity Recognition (NER) is ideal—it automatically spots geographic entities (tagged as GPE, which covers countries, cities, and regions).
- Install spaCy and the English language model:
pip3 install spacy python3 -m spacy download en_core_web_sm - Example code:
If you want to separate countries from cities, you can cross-check the results with pycountry's database to filter out non-country entries.import spacy # Load the pre-trained English NER model nlp = spacy.load("en_core_web_sm") def extract_locations(text): doc = nlp(text) # Extract all entities labeled as geographic locations locations = [ent.text for ent in doc.ents if ent.label_ == "GPE"] return locations # Test cases print(extract_locations("I live in Spain")) # Output: ['Spain'] print(extract_locations("United States (New York), United Kingdom (London)")) # Output: ['United States', 'New York', 'United Kingdom', 'London']
If your text has vague or incomplete location references, geopy can use geocoding services (like OpenStreetMap's Nominatim) to parse and validate locations. Note: This requires an internet connection.
- Install geopy:
pip3 install geopy - Example code:
from geopy.geocoders import Nominatim def extract_country_from_text(text): geolocator = Nominatim(user_agent="location_extractor") location_results = geolocator.geocode(text, exactly_one=False) countries = [] if location_results: for loc in location_results: # Pull country name from the address components country = loc.raw.get('address', {}).get('country') if country and country not in countries: countries.append(country) return countries # Test cases print(extract_country_from_text("I live in Spain")) # Output: ['España'] (map to English with pycountry if needed) print(extract_country_from_text("United States (New York), United Kingdom (London)")) # Output: ['United States of America', 'United Kingdom']
One critical reminder: Since you have both Python 2 and 3 installed, always use pip3 for installing packages and python3 for running scripts—this ensures everything uses your Python 3 environment instead of Python 2.
内容的提问来源于stack exchange,提问作者Markus

