如何将属性文本转换为UTF8字符串?特殊字符转换技术咨询
Hey there! Let's break down your two attributed text questions one by one:
The approach here depends on whether you want just the plain text content or need to preserve formatting attributes in the output string. Here are the two common scenarios:
Scenario 1: Extract Plain Text (No Attributes)
If you only care about raw text and don't need to keep formatting like fonts, colors, or styles, it's straightforward:
- Grab the plain
stringproperty from yourNSAttributedString/NSMutableAttributedString - Convert it to a UTF-8 encoded data object, then back to a UTF-8 string.
// Assume you have an NSAttributedString *attributedText NSString *plainText = attributedText.string; NSData *utf8Data = [plainText dataUsingEncoding:NSUTF8StringEncoding]; NSString *utf8String = [[NSString alloc] initWithData:utf8Data encoding:NSUTF8StringEncoding];
Scenario 2: Preserve Attributes (Serialize as HTML/RTF)
If you need to retain formatting, you can serialize the attributed text to HTML or RTF (both UTF-8 compatible formats), then convert that to a UTF-8 string.
HTML Serialization Example
NSAttributedString *attributedText = /* Your attributed text */; NSData *htmlData = [attributedText dataFromRange:NSMakeRange(0, attributedText.length) documentAttributes:@{NSDocumentTypeDocumentAttribute: NSHTMLTextDocumentType} error:nil]; NSString *utf8HtmlString = [[NSString alloc] initWithData:htmlData encoding:NSUTF8StringEncoding];
RTF Serialization Example
NSData *rtfData = [attributedText dataFromRange:NSMakeRange(0, attributedText.length) documentAttributes:@{NSDocumentTypeDocumentAttribute: NSRTFTextDocumentType} error:nil]; NSString *utf8RtfString = [[NSString alloc] initWithData:rtfData encoding:NSUTF8StringEncoding];
Let's unpack why your code turns \Ud835\Udc13\Ud835\Udc1e\Ud835\Udc2c\Ud835\Udc2d into properly rendered "Test" (or more precisely, mathematical bold capital letters TEST):
First, decode that Unicode sequence:
\Ud835\Udc13= U+1D413 (Mathematical Bold Capital T)\Ud835\Udc1e= U+1D41E (Mathematical Bold Capital E)\Ud835\Udc2c= U+1D42C (Mathematical Bold Capital S)\Ud835\Udc2d= U+1D42D (Mathematical Bold Capital T)
These are UTF-16 surrogate pairs — they represent Unicode characters outside the Basic Multilingual Plane (BMP), which require two 16-bit code units to encode. Here's how your code processes them:
String to UTF-8 Data:
[self.message dataUsingEncoding:NSUTF8StringEncoding]converts the UTF-16 string (with surrogate pairs) into a valid UTF-8 byte stream. UTF-8 handles non-BMP characters by encoding them into 4-byte sequences, so this step correctly preserves all character data.NSHTMLTextDocumentType Parsing:
When initializingNSMutableAttributedStringwith theNSHTMLTextDocumentTypeoption, the system uses WebKit's HTML rendering engine under the hood.- Even though your input is raw Unicode characters (not full HTML markup), WebKit treats it as plain text in an implicit HTML document.
- It correctly interprets the UTF-8 encoded surrogate pairs as their corresponding Unicode characters (the bold math letters) and generates an attributed string with the necessary font metadata to render them (assuming the system has a supporting font like San Francisco).
Character Encoding Handling:
TheNSCharacterEncodingDocumentAttribute: NSUTF8StringEncodingparameter explicitly tells the parser the input data is UTF-8 encoded, preventing any byte sequence misinterpretation.
In short, your code leverages WebKit's robust Unicode and HTML handling to convert raw surrogate pair data into a properly rendered attributed string.
内容的提问来源于stack exchange,提问作者Sunny Shah

