如何高效用Pandas实现两个DataFrame的词干匹配关联检索?
如何高效匹配铭文文本中的位置词干对应形容词?
我有两个Pandas DataFrame:一个是约35000条的位置数据,另一个是约550000条的拉丁语铭文数据,简化结构如下:
locations DataFrame
locations period name long lat latlong [...] stem 0 R Roma 50.4 1.23 50.4,1.23 [...] Rom 1 GR Londinium 52.2 3.45 52.2,3.45 [...] Londini, Londin 2 RL Colonia Agrippina 54.0 6.78 54.0,6.78 [...] Coloni, Colon, Agrippin 3 G Vindobona 55.8 9.01 55.8,9.01 [...] Vindobon 4 R Lutetia Parisiorum 57.6 -3.21 57.6,-3.21 [...] Luteti, Lutet, Parisior [...]
inscriptions DataFrame
inscriptions id text findspot [...] comment 0 a0 some text Brindisi [...] n/a 1 a1 some different text Bath [...] viri 2 a2 text with name Bern [...] tria nomina 3 a3 text with Romensis Berkeley [...] mulieris 4 a4 long text Bonn [...] n/a [...]
我的需求是找出所有文本中包含位置名称形容词形式的铭文。目前的实现是用df.iterrows()遍历locations的35000行,提取每个位置的词干,用正则匹配铭文文本中的对应形式(比如词干Rom可以匹配铭文3里的Romensis,类似用Boston匹配Bostonian这类同根形容词),匹配成功就把铭文行复制到新DataFrame,并添加该位置的经纬度等信息,最终结果示例:
matches id text findspot [...] comment name long lat latlong [...] 0 a3 text with Romensis Berkeley [...] mulieris Rome 50.4 1.23 52.2,3.45 [...] [...]
但这种嵌套遍历效率极低:运行12小时只完成了4%的位置数据处理,预计要耗时数周。我试过np.where(),但没法同时保留位置名称等信息,也没法收集所有同名或近似名称位置的匹配对;df.apply()配合lambda也有类似局限。
请问有没有不用嵌套循环就能实现相同功能的高效方案?
内容的提问来源于stack exchange,提问作者Thomas Leibundgut
相关产品推荐
相关产品推荐

