如何用正则表达式提取文本中所有UUID格式的ID?
Hey there! I see you're stuck trying to pull all those UUIDs from your JSON data—your current regex only grabs one, but we can fix that easily. Let's break down the solutions:
Option 1: Use a Precise Regex with Global Matching
Your original regex is too vague ((\d*\w*-\d*\w*-\d*\w*-\d*\w*-\d*\w*)) and missing the global match flag, which is why it only returns the first UUID. Instead, use a regex tailored to UUID's standard 8-4-4-4-12 hex format:
[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}
Add flags to make it work globally and ignore case (since UUIDs might use uppercase letters):
g(global): Matches all occurrences, not just the firsti(ignore case): Captures both uppercase and lowercase hex characters
Example Implementation (JavaScript)
If you're working in a browser or Node.js:
// Your JSON string (replace ... with full content if possible) const jsonContent = `[ { "id": "059ffad4-1f6f-4221-82bb-1a56ca709e28", "contentType": "tv", "Logo": null, "Large": "", "Medium": "", "Thumbnail": " https://imagesiwantnow.blob.core.windows.net/iwantnow/ott/episodes/2018-01-16/20141027-forevermore-289x163.jpg ", "title": "001 - Forevermore - Oct 27, 2014" }, { "id": "dfb35a01-865d-4203-a31c-ff97a6a1fa51", "contentType": "tv", "Logo": null, "Large": "", "Medium": "", ... ]`; // Extract all UUIDs const allUuids = jsonContent.match(/[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}/gi); // Add to a listbox (example for HTML) const listbox = document.getElementById('yourListboxID'); allUuids.forEach(uuid => { const option = document.createElement('option'); option.textContent = uuid; option.value = uuid; listbox.appendChild(option); });
Example Implementation (Python)
If you're using Python:
import re # Your JSON content as a string json_content = '''[ { "id": "059ffad4-1f6f-4221-82bb-1a56ca709e28", "contentType": "tv", "Logo": null, "Large": "", "Medium": "", "Thumbnail": " https://imagesiwantnow.blob.core.windows.net/iwantnow/ott/episodes/2018-01-16/20141027-forevermore-289x163.jpg ", "title": "001 - Forevermore - Oct 27, 2014" }, { "id": "dfb35a01-865d-4203-a31c-ff97a6a1fa51", "contentType": "tv", "Logo": null, "Large": "", "Medium": "", ... ]''' # Extract all UUIDs all_uuids = re.findall(r'[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}', json_content, re.IGNORECASE) # Add to a Tkinter listbox from tkinter import Tk, Listbox root = Tk() listbox = Listbox(root) listbox.pack(padx=10, pady=10) for uuid in all_uuids: listbox.insert('end', uuid) root.mainloop()
Option 2: Parse the JSON Directly (More Reliable)
If your JSON is complete (no ... truncation), parsing it directly is a safer approach—regex can break if the JSON structure changes.
Example (JavaScript)
// Parse the JSON (ensure it's valid, remove ... first) const data = JSON.parse(jsonContent); const allUuids = data.map(item => item.id); // Add to listbox same as before
Example (Python)
import json # Parse valid JSON (remove ... first) data = json.loads(json_content) all_uuids = [item['id'] for item in data] # Add to listbox same as before
Why Your Original Regex Failed
Your regex (\d*\w*-\d*\w*-\d*\w*-\d*\w*-\d*\w*) is overly broad (it matches any word/number combinations separated by dashes) and doesn't include a global match flag. Without g (in JS) or using findall (in Python), it stops after the first match.
内容的提问来源于stack exchange,提问作者zackmark29

