如何过滤字符串中的不可见字符?PHP场景技术求助
Hey there, let's tackle those pesky invisible characters messing with your string processing! I've run into this exact issue before when copying web content—those hidden Unicode control chars are total troublemakers. Here's how to fix it step by step:
Before jumping into fixes, it helps to know exactly what you're dealing with. Add a quick debug snippet to print the Unicode value of the first character of each word—this will reveal if it's a zero-width space (U+200B), non-breaking space (U+00A0), or another control character:
foreach (explode(' ', $this->attributes['Subject_name']) as $word) { $firstChar = mb_substr($word, 0, 1, 'UTF-8'); echo "Word: '$word' | First char Unicode: U+" . strtoupper(dechex(mb_ord($firstChar))) . "\n"; }
The most reliable way to strip these hidden chars is using a Unicode-aware regex to target non-printable control/format characters. Combine this with replacing non-breaking spaces (a common culprit from web text) with regular spaces:
// Clean the original string first $str = $this->attributes['Subject_name']; // Remove all Unicode control/format/invisible chars $str = preg_replace('/[\p{Cc}\p{Cf}\p{Cs}\p{Cn}]/u', '', $str); // Replace non-breaking spaces with regular spaces $str = str_replace("\xc2\xa0", ' ', $str); // Trim and collapse extra spaces to avoid empty "words" $str = preg_replace('/\s+/', ' ', trim($str));
Now that the string is clean, update your code to use mb_substr instead of substr (critical for multi-byte UTF-8 characters) and make your exclude list case-insensitive:
public function getPrefixAttribute() { $str = $this->attributes['Subject_name']; // Clean invisible characters first $str = preg_replace('/[\p{Cc}\p{Cf}\p{Cs}\p{Cn}]/u', '', $str); $str = str_replace("\xc2\xa0", ' ', $str); $str = preg_replace('/\s+/', ' ', trim($str)); $exclude = array('And', 'of', 'in'); // Normalize exclude list to lowercase for case-insensitive checks $lowerExclude = array_map('strtolower', $exclude); $ret = ''; foreach (explode(' ', $str) as $word) { $word = trim($word); if (in_array(strtolower($word), $lowerExclude)) { continue; } // Use mb_substr to safely get first character (supports UTF-8) $firstChar = mb_substr($word, 0, 1, 'UTF-8'); $ret .= strtoupper($firstChar); } return $ret; }
To avoid dealing with this in the future, add a cleanup step before saving data to the database. Use your model's saving hook:
protected static function boot() { parent::boot(); static::saving(function ($model) { if (isset($model->Subject_name)) { $model->Subject_name = preg_replace('/[\p{Cc}\p{Cf}\p{Cs}\p{Cn}]/u', '', $model->Subject_name); $model->Subject_name = str_replace("\xc2\xa0", ' ', $model->Subject_name); $model->Subject_name = preg_replace('/\s+/', ' ', trim($model->Subject_name)); } }); }
内容的提问来源于stack exchange,提问作者Yogesh.galav

