You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从大字符串中检索匹配行并生成DataFrame

从文本提取指定子串行并生成带翻译的DataFrame

1. 依赖安装

首先确保安装pandas库,用于生成结构化DataFrame:

pip install pandas

2. 核心实现代码

import pandas as pd

# 定义需要检索的目标英文子串
target_substrings = [
    "democracy’s day",
    "America has been tested",
    "celebrate the triumph",
    "democracy is precious"
]

# 替换为你的实际大文本内容
raw_text = """
Today is democracy’s day, a time to honor our nation's values.
Through every crisis, America has been tested and emerged stronger.
We gather here to celebrate the triumph of hope over despair.
Let us remember that democracy is precious and must be defended.
This line doesn't match any target substring.
Another line without the target phrases.
"""

# 分割文本为单行,过滤空行并清理首尾空格
lines = [line.strip() for line in raw_text.split('\n') if line.strip()]

# 匹配并提取目标行,同时添加对应中文翻译
matched_records = []
translation_map = {
    "democracy’s day": "民主日",
    "America has been tested": "美国历经考验",
    "celebrate the triumph": "庆祝胜利",
    "democracy is precious": "民主弥足珍贵"
}

for line in lines:
    for substr in target_substrings:
        if substr in line:
            matched_records.append({
                "英文原句": line,
                "匹配子串": substr,
                "中文翻译": translation_map[substr]
            })
            break  # 避免同一行匹配多个子串重复录入

# 生成DataFrame
result_df = pd.DataFrame(matched_records)

# 可选择导出为CSV文件:result_df.to_csv('matched_lines.csv', index=False, encoding='utf-8-sig')
print(result_df)

3. 关键逻辑说明

  • 文本预处理:将大文本按换行符分割为单行,过滤空行并去除每行首尾空格,排除无效数据;
  • 匹配与翻译:遍历每行文本,检查是否包含目标子串,匹配成功后同步记录原句、匹配子串及对应中文翻译,break语句确保同一行仅被记录一次;
  • 结构化输出:将匹配结果转换为DataFrame,支持直接打印查看或导出为CSV等格式。

内容的提问来源于stack exchange,提问作者san1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 12:01:02