基于stringr环视功能批量提取文献摘要中的植物种名
问题
我有一个包含约20万条观测数据的DataFrame,其中某列存储着科学期刊的摘要,需要从中提取植物种名。目前已从摘要中提取出属名,种名通常是「属名+种加词」的格式(比如Genus species),本来可以用正则环视功能匹配,但属名多达数千个,手动构建pattern = Malus|Gentiana|Acer|Quercus这类正则模式根本不现实。我想了解是否有办法(比如函数)能自动将DataFrame单列中的属名代入环视规则完成匹配,同时支持处理单篇摘要提及多个属名的情况。
示例场景
示例1
摘要内容:
axillary bud cultures were initiated from 3 types of nodal explants of lagerstroemia parviflora. the cultures derived from explants of seedlings, terminal twigs of a 50-year-old tree and basal-sprouts of another 50-year-old tree showed significant variation in responses at establishment, shoot proliferation and rooting stages. all the 3 types of explants exuded phenolic substances from their cut ends. the exudation was checked by suspending them in a solution of 25 mu m pvp 40 and 522.5 mu m citric acid; and by the addition of 100 mu m pvp 40 and 522.5 mu m citric acid in establishment medium. leaching continued upto rooting stage, therefore, pvp 40 and citric acid were added in ms medium used for successive transfer and rooting of microshoots. seedling and basal-sprout explants placed on ms medium with 0.44 mu m ba showed maximum shoot lengths 1.45 cm +/- 0.13 and 1.16 cm +/- 0.22, respectively. tree explants exhibited best axillary shoot elongation (0.8 cm +/- 0.07) on ms medium without plant growth regulators. the cultures derived from seedling and basal-sprout explants could be successfully maintained upto 6th successive transfers, whereas, those derived from tree explants died after 3rd transfer. microshoots obtained from seedling and basal-sprout explants showed 10% rooting on ms medium supplemented with 4.9 mu m iba.期望提取结果:
lagerstroemia parviflora
示例2
摘要内容:
the influence of headspace ethylene on anthocyanin, anthocyanidin, and carotenoid accumulation was studied in suspension cultures of vaccinium pahalae. exogenous application of ethrel (an ethylene-releasing compound) significantly reduced growth and secondary metabolite production, whereas incorporation of 5.0 or 10.0 mg l(-1) clcl(2) or nicl(2) effectively reduced ethylene accumulation and improved product accumulation, but agno(3) was toxic to cells. this study showed an overall negative impact of increased ethylene levels in the vessel headspace on phytochemical production in ohelo cell cultures.期望提取结果:
vaccinium pahalae
多属名示例
摘要内容:
in two turfgrass species, festuca arundinacea schreb, (tall fescue) and zoysia japonica steud, (zoysiagrass), regeneration culture systems using two types of bioreactors were developed, regenerants of tall fescue and zoysiagrass were efficiently produced by using an aeration-agitation type bioreactor and a rotating drum type bioreactor, respectively, the regenerants of each species were harvested from the bioreactors and cultivated in vitro during the preparation stage either on a 1/4 strength ng gellan gum (4 g l(-1)) medium without sucrose or with 30 g l(-1) sucrose, and under co2 concentration of 0.4 or 50 mmol mol(-1), a photoperiod of 24 h per day and a photosynthetic photon flux density of 125 mu mol m(-2) s(-1). the shoot and root lengths and shoot and root dry weights of tall fescue regenerants and the root dry weight of zoysiagrass regenerants were greater on the medium with sucrose than those on the medium without sucrose, regardless of the co2 concentration, the survival percentage, shoot number and shoot length of zoysiagrass regenerants growing on the medium without sucrose under 50 mmol mol(-1)of co2 were the highest among all the treatments.期望提取结果:
festuca arundinacea schreb和zoysia japonica steud
解决方案
核心思路
- 从DataFrame的属名列中提取所有唯一属名,构建正则匹配的备选列表
- 编写正则表达式,自动嵌入属名列表,匹配「属名 + 后续种加词/命名人信息」的完整种名格式
- 用Pandas批量处理摘要列,提取所有匹配的种名,支持单篇摘要多属名的情况
代码实现(Python)
假设你的DataFrame名为df,摘要列是abstract,已提取的属名列是genus:
import pandas as pd import re # 1. 提取所有唯一属名,转义后构建正则模式字符串 unique_genera = df['genus'].dropna().unique() genus_pattern = '|'.join(re.escape(genus) for genus in unique_genera) # 2. 编写完整正则:匹配属名开头的完整种名(含种加词、命名人) # 负向环视确保属名前不是单词字符,避免错误匹配;忽略大小写适应摘要格式 species_regex = re.compile( r'(?<!\w)({}(\s+[a-zA-Z.,-]+)+)'.format(genus_pattern), re.IGNORECASE ) # 3. 定义提取函数:处理单条摘要,返回所有匹配的唯一种名 def extract_species_from_abstract(abstract): if pd.isna(abstract): return [] # 提取所有完整匹配项,去除末尾逗号,去重后返回 matches = species_regex.findall(abstract) unique_species = list(set([match[0].rstrip(',') for match in matches])) return unique_species # 4. 批量应用到DataFrame df['extracted_species'] = df['abstract'].apply(extract_species_from_abstract)
关键细节说明
- 属名转义:用
re.escape()处理每个属名,避免属名中包含.、-等正则特殊字符导致匹配失效 - 正则规则:
(?<!\w):负向环视,确保属名前面不是单词字符(比如不会把Xylosma lagerstroemia中的lagerstroemia误判为属名)(\s+[a-zA-Z.,-]+)+:匹配属名后跟随的种加词、命名人缩写(支持含连字符、点号、逗号的格式)re.IGNORECASE:忽略大小写,适配摘要中属名首字母大写或全小写的情况
- 多匹配处理:用
findall提取所有符合规则的种名,通过set去重后返回列表,自然支持单篇摘要多个种名的场景
灵活调整建议
- 如果不需要提取命名人信息,可将正则简化为:
r'(?<!\w)({})\s+[a-zA-Z]+'.format(genus_pattern) - 如果种名包含数字或其他特殊字符,可扩展正则中的字符范围,比如把
[a-zA-Z.,-]改成[a-zA-Z.,-0-9]
内容的提问来源于stack exchange,提问作者Melissa Duda
相关产品推荐
相关产品推荐

