多词州名匹配优化:Python实现美国大学城字典的解决方案
Fixing University Towns Data Parsing (No Regex Needed)
Hey there! Let's tackle those two issues you're facing with your code—no regular expressions required, just some targeted string operations. The core problems were splitting lines by space too early (which broke multi-word state and city names) and not properly identifying/processing state vs. town entries.
The Original Issues Recap
- Multi-word states (like North Carolina) weren't being recognized because your code only took the first word of each line.
- City names with spaces (like Isla Vista) were getting split incorrectly for the same reason.
Revised Code
import pandas as pd def get_list_of_university_towns(): ''' Returns a DataFrame of towns and the states they are in from the university_towns.txt list. The format of the DataFrame should be: DataFrame( [ ["Michigan", "Ann Arbor"], ["Michigan", "Yipsilanti"] ], columns=["State", "RegionName"] ) The following cleaning needs to be done: 1. For "State", removing characters from "[" to the end. 2. For "RegionName", when applicable, removing every character from " (" to the end. 3. Depending on how you read the data, you may need to remove newline character '\n'. ''' # Read file, strip whitespace from each line, and skip empty lines with open("university_towns.txt") as fhandle: lines = [line.strip() for line in fhandle if line.strip()] current_state = None town_records = [] for line in lines: # Check if this line is a state (ends with [edit]) if line.endswith("[edit]"): # Extract state name by cutting off everything from "[" onwards current_state = line[:line.find("[")].strip() else: # Clean up town name: remove everything starting from " (" if " (" in line: clean_town = line[:line.find(" (")].strip() else: clean_town = line.strip() # Add the state-town pair to our records list if current_state: town_records.append([current_state, clean_town]) # Convert the list of pairs to the required DataFrame return pd.DataFrame(town_records, columns=["State", "RegionName"])
Key Improvements
- Proper state identification: Instead of splitting lines, we check if a line ends with
[edit]to spot state entries. We then slice the string up to the[to get the full state name (including spaces). - Clean city names: For town lines, we look for the
(character and slice before it to remove university info or citations—keeping the full city name intact. - Safer file handling: Using the
withstatement ensures the file is automatically closed after reading, avoiding resource leaks. - Empty line filtering: We skip any blank lines in the input file to avoid errors.
- Direct DataFrame building: We build a list of
[State, RegionName]pairs directly, which makes converting to the required DataFrame straightforward.
This code will correctly handle multi-word states like North Carolina and cities like Isla Vista, and return the DataFrame in the format you need.
内容的提问来源于stack exchange,提问作者CuriousLearner
相关产品推荐
相关产品推荐

