You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从txt文档提取粗体斜体文本?求适配Mac的R语言脚本解决方案

前置说明

默认你的txt文档中的粗斜体采用Markdown标准标记规则:粗体用**内容**包裹,斜体用*内容*包裹。如果是其他标记规则,修改对应正则表达式即可。

R语言实现方案(Mac原生适配)

依赖安装

打开Mac终端,依次执行以下命令安装环境和依赖:

  • 安装R环境(已安装可跳过):brew install r
  • 安装字符串处理包:终端执行Rscript -e 'install.packages("stringr")'

脚本代码

将以下代码保存为extract_formatted_text.R,修改开头的目录参数为你自己的路径:

# 加载依赖
library(stringr)

# 自定义参数 请根据实际情况修改
txt_dir <- "~/Documents/txt_files" # 存放txt文件的文件夹路径
output_path <- "~/Documents/extracted_result.txt" # 输出结果的文件路径

# 读取所有txt文件
txt_files <- list.files(path = txt_dir, pattern = "\\.txt$", full.names = TRUE, recursive = TRUE)
all_extracted <- c()

for (file in txt_files) {
  # 跳过输出文件本身 避免重复读取
  if (file == output_path) next
  # 读取文件内容 出现乱码可修改encoding参数适配
  content <- readLines(file, encoding = "UTF-8", warn = FALSE)
  content <- paste(content, collapse = "\n")
  
  # 匹配粗体内容 **xxx**
  bold_matches <- str_match_all(content, "\\*\\*(.*?)\\*\\*")[[1]][,2]
  # 匹配斜体内容 *xxx* 排除粗体标记干扰
  content_no_bold <- str_replace_all(content, "\\*\\*.*?\\*\\*", "")
  italic_matches <- str_match_all(content_no_bold, "\\*(.*?)\\*")[[1]][,2]
  
  # 合并结果
  all_extracted <- c(all_extracted, bold_matches, italic_matches)
}

# 去重 去除空值
all_extracted <- unique(na.omit(all_extracted))
all_extracted <- all_extracted[all_extracted != ""]

# 导出结果
writeLines(all_extracted, con = output_path, encoding = "UTF-8")
cat("提取完成,结果已保存到:", output_path, "\n")

运行方法

打开终端,进入脚本存放的目录,执行:
Rscript extract_formatted_text.R
运行后即可在你设置的输出路径拿到提取结果。


备选方案:Python实现(大部分Mac机型自带环境无需额外安装)

将以下代码保存为extract_format.py,修改路径后终端执行python3 extract_format.py即可:

import os
import re

# 自定义参数
txt_dir = "~/Documents/txt_files"
output_path = "~/Documents/extracted_result.txt"

# 路径格式化
txt_dir = os.path.expanduser(txt_dir)
output_path = os.path.expanduser(output_path)

extracted = set()

for root, _, files in os.walk(txt_dir):
    for file in files:
        if not file.endswith(".txt"):
            continue
        file_path = os.path.join(root, file)
        if file_path == output_path:
            continue
        with open(file_path, "r", encoding="utf-8") as f:
            content = f.read()
        # 匹配粗体
        bold = re.findall(r"\*\*(.*?)\*\*", content)
        extracted.update(bold)
        # 去掉粗体标记后匹配斜体
        content_no_bold = re.sub(r"\*\*.*?\*\*", "", content)
        italic = re.findall(r"\*(.*?)\*", content_no_bold)
        extracted.update(italic)

# 去除空值
extracted = [i for i in extracted if i.strip()]

with open(output_path, "w", encoding="utf-8") as f:
    f.write("\n".join(extracted))
print("提取完成")

内容的提问来源于stack exchange,提问作者GIS_newbie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 04:06:05