You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Pandas按出现位置提取指定条件的文本

解决步骤与代码实现

没问题,我来帮你搞定这个需求,下面是具体的代码和解释:

首先,我们先把示例数据转换成Pandas DataFrame:

import pandas as pd

# 构建示例数据
data = {
    '页码': [1,1,1,2,2,2,2,2,2,2,2,2,3],
    '文本': ['they','ate','apples','and','then','their','stomach','felt','slightly','painful','given','the','fruit'],
    '字体': [0,0,0,0,1,0,0,1,0,1,0,1,0]
}
df = pd.DataFrame(data)

接下来完成两个核心操作:

  1. 查找单词'then'首次出现的页码
# 找到'then'首次出现的页码
target_page = df[df['文本'] == 'then']['页码'].iloc[0]
  1. 在目标页码下,提取'then'之后出现的、font列值为1的前三个文本
# 筛选目标页码的所有行
page_data = df[df['页码'] == target_page]
# 找到'then'所在的索引位置
then_index = page_data[page_data['文本'] == 'then'].index[0]
# 取'then'之后的行,筛选font=1的条目,再取前三个文本转成列表
result = page_data.loc[then_index+1:][page_data['字体'] == 1]['文本'].head(3).tolist()

print(result)  # 输出: ['felt', 'painful', 'the']

这里的逻辑很直观:先锁定目标页码,再定位到'then'的位置,只取它之后的内容,接着筛选出字体为1的条目,最后提取前三个文本就得到了想要的结果。

内容的提问来源于stack exchange,提问作者user17312322

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 10:25:43