如何用正则表达式实现多端JS代码块的拆分生成?
Got it, let's break this down step by step—you want to split JS code into Web and mobile versions using custom markers, avoid regex headaches, and fix nested block issues. Here's how to do it properly:
1. Pick Low-Conflict, Easy-to-Match Markers
First, let's ditch the {*Phone ... */Phone} style—it's too close to JS syntax (like template literals or object literals) and risks accidental matches. Instead, use JS-comment-compatible markers with unique separators that'll never show up in regular code:
- Web code blocks: Start with
/*::WEB-BEGIN::*/, end with/*::WEB-END::*/ - Mobile code blocks: Start with
/*::PHONE-BEGIN::*/, end with/*::PHONE-END::*/
Why this works:
- These are valid JS comments, so your raw source code will run without breaking even before splitting
- The
::separators and explicit BEGIN/END make the markers unique—no chance of them appearing in business code by accident - Regex matching becomes trivial because the boundaries are crystal clear
2. Regex Implementation for Non-Nested Blocks
If you can keep your code blocks flat (no nested same-type blocks), simple regex will do the trick. Here's how to extract and split the code:
Extract & Split Function
function splitFlatCode(sourceCode) { // Regex patterns for Web and Phone blocks const webBlockRegex = /\/\*::WEB-BEGIN::\*\/([\s\S]*?)\/\*::WEB-END::\*\//g; const phoneBlockRegex = /\/\*::PHONE-BEGIN::\*\/([\s\S]*?)\/\*::PHONE-END::\*\//g; // Generate Web code: keep non-Phone blocks + Web block content const webCode = sourceCode .replace(phoneBlockRegex, '') // Remove all Phone blocks .replace(webBlockRegex, '$1'); // Replace Web blocks with their inner content // Generate Phone code: keep non-Web blocks + Phone block content const phoneCode = sourceCode .replace(webBlockRegex, '') // Remove all Web blocks .replace(phoneBlockRegex, '$1'); // Replace Phone blocks with their inner content return { webCode: webCode.trim(), phoneCode: phoneCode.trim() }; }
Regex Breakdown
\/\*::WEB-BEGIN::\*\/: Escapes the/*and*/comment syntax to match the start marker exactly([\s\S]*?): Non-greedily matches any character (including newlines, thanks to[\s\S]) inside the block\/\*::WEB-END::\*\//: Matches the end marker exactly
3. Fixing Nested Block Issues
Plain regex can't handle nested blocks (like a Web block inside another Web block) because JavaScript regex doesn't support native recursion. Instead, use a stack-based approach to track nested block levels:
Nested Block Split Function
function splitNestedCode(sourceCode) { const webCodeBuffer = []; const phoneCodeBuffer = []; const blockStack = []; // Tracks current nested block types: 'WEB' or 'PHONE' // Regex to match all BEGIN/END markers const markerRegex = /\/\*::(WEB|PHONE)-(BEGIN|END)::\*\//g; let lastPosition = 0; let match; while ((match = markerRegex.exec(sourceCode)) !== null) { const [fullMarker, blockType, action] = match; const textBetweenMarkers = sourceCode.slice(lastPosition, match.index); // Add the text between markers to the correct buffers if (blockStack.length === 0) { // Outside any block: add to both buffers webCodeBuffer.push(textBetweenMarkers); phoneCodeBuffer.push(textBetweenMarkers); } else { const currentTopBlock = blockStack[blockStack.length - 1]; if (currentTopBlock === 'WEB') { webCodeBuffer.push(textBetweenMarkers); } else if (currentTopBlock === 'PHONE') { phoneCodeBuffer.push(textBetweenMarkers); } } // Update the block stack based on BEGIN/END action if (action === 'BEGIN') { blockStack.push(blockType); } else { // Find and remove the matching BEGIN from the stack const matchingBeginIndex = blockStack.lastIndexOf(blockType); if (matchingBeginIndex !== -1) { blockStack.splice(matchingBeginIndex); } } lastPosition = markerRegex.lastIndex; } // Add the final chunk of text after the last marker const finalText = sourceCode.slice(lastPosition); if (blockStack.length === 0) { webCodeBuffer.push(finalText); phoneCodeBuffer.push(finalText); } else { const currentTopBlock = blockStack[blockStack.length - 1]; if (currentTopBlock === 'WEB') { webCodeBuffer.push(finalText); } else if (currentTopBlock === 'PHONE') { phoneCodeBuffer.push(finalText); } } return { webCode: webCodeBuffer.join('').trim(), phoneCode: phoneCodeBuffer.join('').trim() }; }
How This Works
- The
blockStacktracks which type of block we're currently inside (even nested ones) - For each marker, we add the text between markers to the correct buffer based on the current stack state
- When we hit an END marker, we pop the matching BEGIN from the stack to properly close nested blocks
4. Pro Tips to Avoid Headaches
- Avoid nesting if possible: Flat code blocks are easier to maintain and require simpler regex. Only use nesting if it's absolutely necessary.
- Validate markers first: Add a quick check to ensure every BEGIN marker has a matching END marker—this prevents broken splits from typos.
- Integrate with your build pipeline: Turn this logic into a Webpack/Vite plugin or CLI tool so you don't have to run it manually every time.
内容的提问来源于stack exchange,提问作者DDave

