You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用正则表达式实现文本指定词汇替换的技术咨询

Alright, let's work through your regex matching and replacement challenge. First, let's unpack the behavior of your current pattern, then show you how to implement the exact replacement you need.

1. Quick Analysis of Your Regex

Your pattern \b(Journal|Year|Page|DOI|Roche-Link)\b is mostly correct for matching the target terms, but let's clarify one key detail:

  • \b (word boundary) works perfectly for Roche-Link because it sits between word characters (the e in Roche, L in Link) and non-word characters (the - and surrounding spaces). So it will correctly match the full phrase when it's isolated by spaces or at the start/end of the string.
  • The real issue isn't the matching itself—it's that you need each matched term to map to a unique replacement string, which requires dynamic replacement logic instead of a static replace call.

2. Implementation Examples (Common Languages)

Here's how to set up the replacement using a mapping dictionary and regex substitution with a callback function:

Python

import re

input_text = "Journal Page Year DOI Roche-Link"
# Define your replacement mappings clearly
term_mapping = {
    "Journal": "Publication Title",
    "Page": "Pagination",
    "Year": "Publication Date",
    "DOI": "Digital Object Identifier",
    "Roche-Link": "Roche Link"
}

# Use re.sub with a lambda to look up the replacement for each match
output_text = re.sub(r'\b(Journal|Year|Page|DOI|Roche-Link)\b', lambda match: term_mapping[match.group()], input_text)
print(output_text)
# Output: Publication Title Pagination Publication Date Digital Object Identifier Roche Link

JavaScript

const inputText = "Journal Page Year DOI Roche-Link";
const termMapping = {
  "Journal": "Publication Title",
  "Page": "Pagination",
  "Year": "Publication Date",
  "DOI": "Digital Object Identifier",
  "Roche-Link": "Roche Link"
};

// Use String.replace with a callback to fetch the correct replacement
const outputText = inputText.replace(/\b(Journal|Year|Page|DOI|Roche-Link)\b/g, (matchedTerm) => termMapping[matchedTerm]);
console.log(outputText);
// Output: Publication Title Pagination Publication Date Digital Object Identifier Roche Link

3. Optional Adjustments

  • Case Insensitivity: If your input might have lowercase or mixed-case versions of the terms (e.g., journal, doi), add the case-insensitive flag and adjust the mapping to handle variations:
    # Python example with case insensitivity
    output_text = re.sub(r'\b(Journal|Year|Page|DOI|Roche-Link)\b', lambda m: term_mapping[m.group().capitalize()], input_text, flags=re.IGNORECASE)
    
  • Flexible Boundaries: If your terms might appear next to non-word characters (like parentheses or hyphens) that \b doesn't recognize, replace \b with lookaround assertions for better control:
    (?<!\w)(Journal|Year|Page|DOI|Roche-Link)(?!\w)
    
    This ensures the term isn't preceded or followed by any word character (letters, numbers, underscores), regardless of other surrounding punctuation.

内容的提问来源于stack exchange,提问作者MokiNex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:59:01