You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效用Pandas实现两个DataFrame的词干匹配关联检索?

如何高效匹配铭文文本中的位置词干对应形容词?

我有两个Pandas DataFrame:一个是约35000条的位置数据,另一个是约550000条的拉丁语铭文数据,简化结构如下:

locations DataFrame

locations
     period               name  long    lat     latlong  [...]                     stem
0        R                Roma  50.4   1.23   50.4,1.23  [...]                      Rom
1       GR           Londinium  52.2   3.45   52.2,3.45  [...]          Londini, Londin
2       RL   Colonia Agrippina  54.0   6.78   54.0,6.78  [...]  Coloni, Colon, Agrippin
3        G           Vindobona  55.8   9.01   55.8,9.01  [...]                 Vindobon
4        R  Lutetia Parisiorum  57.6  -3.21  57.6,-3.21  [...]  Luteti, Lutet, Parisior
[...]

inscriptions DataFrame

inscriptions
     id                 text  findspot  [...]      comment
0    a0            some text  Brindisi  [...]          n/a
1    a1  some different text  Bath      [...]         viri
2    a2       text with name  Bern      [...]  tria nomina
3    a3   text with Romensis  Berkeley  [...]     mulieris
4    a4            long text  Bonn      [...]          n/a
[...]

我的需求是找出所有文本中包含位置名称形容词形式的铭文。目前的实现是用df.iterrows()遍历locations的35000行,提取每个位置的词干,用正则匹配铭文文本中的对应形式(比如词干Rom可以匹配铭文3里的Romensis,类似用Boston匹配Bostonian这类同根形容词),匹配成功就把铭文行复制到新DataFrame,并添加该位置的经纬度等信息,最终结果示例:

matches
     id                 text  findspot  [...]      comment  name  long    lat     latlong  [...]
0    a3   text with Romensis  Berkeley  [...]     mulieris  Rome  50.4   1.23   52.2,3.45  [...]
[...]

但这种嵌套遍历效率极低:运行12小时只完成了4%的位置数据处理,预计要耗时数周。我试过np.where(),但没法同时保留位置名称等信息,也没法收集所有同名或近似名称位置的匹配对;df.apply()配合lambda也有类似局限。

请问有没有不用嵌套循环就能实现相同功能的高效方案?


内容的提问来源于stack exchange,提问作者Thomas Leibundgut

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 09:25:26