You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用JavaScript识别字符串中未大写的专有名词与缩略词?

Practical Strategies for Detecting & Capitalizing Proper Nouns/Acronyms in JavaScript

Hey there! Let's dive into actionable, scalable approaches for your text formatting tool. Since you already know we can't cover every single proper noun or acronym out there, we'll focus on balancing accuracy, ease of implementation, and adaptability for your typical 200-word input.


1. Maintain Core Word Banks (Dual Categories)

Start with two focused maps: one for acronyms (direct lowercase-to-uppercase mappings) and another for full proper noun phrases. This gives you reliable, deterministic results for common terms, and works great for your input size since Map lookups are fast.

Example Code:

// Default acronym mappings (expand as needed)
const baseAcronyms = new Map([
  ['usa', 'USA'],
  ['uk', 'UK'],
  ['nasa', 'NASA'],
  ['fbi', 'FBI']
]);

// Default proper noun mappings (include full phrases)
const baseProperNouns = new Map([
  ['united states of america', 'United States of America'],
  ['new york city', 'New York City'],
  ['international business machines', 'International Business Machines']
]);

function formatText(text) {
  let formatted = text;

  // Handle full phrases first (to avoid partial matches with acronyms)
  baseProperNouns.forEach((corrected, lowercase) => {
    const wordBoundaryRegex = new RegExp(`\\b${lowercase}\\b`, 'gi');
    formatted = formatted.replace(wordBoundaryRegex, corrected);
  });

  // Then process acronyms
  baseAcronyms.forEach((corrected, lowercase) => {
    const wordBoundaryRegex = new RegExp(`\\b${lowercase}\\b`, 'gi');
    formatted = formatted.replace(wordBoundaryRegex, corrected);
  });

  return formatted;
}

Pro Tip: You can seed these maps with public, curated lists (like common country names, global brands, or industry-specific terms) instead of building everything from scratch.


2. Rule-Based Detection for Common Acronym Patterns

Many acronyms follow predictable patterns—think dotted versions (u.s.a) or space-separated single letters (u s a). Add simple regex rules to normalize these into proper uppercase acronyms without needing every term in your word bank.

Example Code:

function normalizeAcronymPatterns(text) {
  // Fix dotted lowercase acronyms (e.g., "u.s.a" → "USA")
  const dottedAcronymRegex = /\b([a-z])\.([a-z])\.([a-z])\b/gi;
  text = text.replace(dottedAcronymRegex, (_, p1, p2, p3) => 
    `${p1.toUpperCase()}${p2.toUpperCase()}${p3.toUpperCase()}`
  );

  // Fix space-separated single letters (e.g., "u s a" → "USA")
  const spacedAcronymRegex = /\b([a-z])\s([a-z])\s([a-z])\b/gi;
  text = text.replace(spacedAcronymRegex, (_, p1, p2, p3) => 
    `${p1.toUpperCase()}${p2.toUpperCase()}${p3.toUpperCase()}`
  );

  return text;
}

Pro Tip: Extend this for 4-letter acronyms if your use case calls for it, but keep rules simple to avoid false positives.


3. Context-Aware Capitalization for Candidate Proper Nouns

For multi-word phrases not in your word bank, use contextual clues to guess which might be proper nouns. For example, phrases following prepositions like in or from, or phrases at the start of a sentence, are likely candidates for title case.

Example Code:

function capitalizeContextualCandidates(text) {
  // Capitalize phrases after common prepositions
  const triggerPrepositions = ['in', 'from', 'to', 'at'];
  triggerPrepositions.forEach(prep => {
    const phraseRegex = new RegExp(`\\b${prep}\\s([a-z]+(\\s[a-z]+)+)\\b`, 'gi');
    text = text.replace(phraseRegex, (match, phrase) => {
      const words = phrase.split(' ');
      const titleCased = words.map(word => {
        // Skip small conjunctions/prepositions in the middle of phrases
        if (word.length <= 2 && ['of', 'the', 'and'].includes(word)) {
          return word;
        }
        return word.charAt(0).toUpperCase() + word.slice(1);
      }).join(' ');
      return `${prep} ${titleCased}`;
    });
  });

  // Capitalize sentence-starting phrases
  const sentenceStartRegex = /^([a-z]+(\s[a-z]+)+)/;
  text = text.replace(sentenceStartRegex, match => 
    match.split(' ').map(word => 
      word.charAt(0).toUpperCase() + word.slice(1)
    ).join(' ')
  );

  return text;
}

Note: This will have occasional false positives, but for 200-word inputs, the tradeoff between automation and accuracy is worth it—plus users can manually adjust if needed.


4. User Customization & Feedback Loops

The best way to fill gaps in your word banks is to let users contribute. Add features for them to define custom acronyms/proper nouns, and even auto-save their corrections to a personal word bank for future use.

Example Code:

// User-customizable storage
let userAcronyms = new Map();
let userProperNouns = new Map();

// Functions to add custom entries
function addUserAcronym(lowercaseInput, correctedValue) {
  userAcronyms.set(lowercaseInput, correctedValue);
}

function addUserProperNoun(lowercaseInput, correctedValue) {
  userProperNouns.set(lowercaseInput, correctedValue);
}

// Updated formatter that merges default and user word banks
function formatTextWithUserCustoms(text) {
  let formatted = text;
  // Merge default and user maps
  const combinedProperNouns = new Map([...baseProperNouns, ...userProperNouns]);
  const combinedAcronyms = new Map([...baseAcronyms, ...userAcronyms]);

  // Process merged maps (same logic as before)
  combinedProperNouns.forEach((corrected, lowercase) => {
    const regex = new RegExp(`\\b${lowercase}\\b`, 'gi');
    formatted = formatted.replace(regex, corrected);
  });

  combinedAcronyms.forEach((corrected, lowercase) => {
    const regex = new RegExp(`\\b${lowercase}\\b`, 'gi');
    formatted = formatted.replace(regex, corrected);
  });

  return formatted;
}

Putting It All Together

Combine these strategies in order of priority to get the best results:

  1. Normalize acronym patterns
  2. Apply user + default proper noun mappings
  3. Apply user + default acronym mappings
  4. Run contextual capitalization for remaining candidates

This workflow ensures you handle known terms first, then clean up common patterns, and finally make educated guesses for unknown phrases.

内容的提问来源于stack exchange,提问作者Roveir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:51:19