如何用GeoText为DataFrame批量提取城市并生成location列?
问题:批量从DataFrame文本列提取城市地理信息报错
现有名为df_test_geo的DataFrame,结构如下:
{'userid': {0: 1, 1: 2, 2: 3, 3: 4, 4: 5, 5: 6, 6: 7}, 'text_string': {0: 'I live in Miami and work in software', 1: 'Chicago, IL', 2: 'Dog Mom in Cincinnati', 3: 'Accountant at @EY/Baltimore', 4: 'World traveler but I call Atlanta home', 5: 'Lover of dogs!', 6: 'Working in Orlando. From Korea.'}}
已导入库:from geotext import GeoText
单独处理单行文本(如places = GeoText(df_test_geo.text_string[0]))可正确返回['Miami'],但直接传入整列(places = GeoText(df_test_geo.text_string))会报错:TypeError: expected string or bytes-like object。
需求:为所有行执行城市提取操作,生成名为location的新列,存储每行提取到的城市(无城市则为空),预期输出顺序为:Miami, Chicago, Cincinnati, Baltimore, Atlanta, '', Orlando
解决方法
GeoText仅支持处理单个字符串对象,无法直接接收Pandas的Series类型,因此需要用apply方法逐行处理text_string列:
代码实现
from geotext import GeoText import pandas as pd # 构造示例DataFrame(已有则可跳过) data = {'userid': {0: 1, 1: 2, 2: 3, 3: 4, 4: 5, 5: 6, 6: 7}, 'text_string': {0: 'I live in Miami and work in software', 1: 'Chicago, IL', 2: 'Dog Mom in Cincinnati', 3: 'Accountant at @EY/Baltimore', 4: 'World traveler but I call Atlanta home', 5: 'Lover of dogs!', 6: 'Working in Orlando. From Korea.'}} df_test_geo = pd.DataFrame(data) # 定义城市提取函数 def extract_city(text): places = GeoText(text) # 提取第一个匹配到的城市,无结果则返回空字符串 return places.cities[0] if places.cities else '' # 生成新列location df_test_geo['location'] = df_test_geo['text_string'].apply(extract_city)
验证结果
执行后df_test_geo['location']的输出为:
0 Miami 1 Chicago 2 Cincinnati 3 Baltimore 4 Atlanta 5 6 Orlando Name: location, dtype: object
完全符合预期要求。
内容的提问来源于stack exchange,提问作者wizkids121
相关产品推荐
相关产品推荐

