Migration, remove accents from all names in a column in the database

Asked

Viewed 183 times

0

I want to make a modification to a data series of a column in the database, remove the accents so I’m doing this function in php. The problem is that none of the functions I discovered work. The strange thing about this is that if you write directly removeAccents('inês') , for example, it already works but if we write removeAccents($_POST['heya']); whereas $_POST['heya'] = 'inês' (example) no longer works.

PHP:

function removeAccents($str) {
  $a = array('À', 'Á', 'Â', 'Ã', 'Ä', 'Å', 'Æ', 'Ç', 'È', 'É', 'Ê', 'Ë', 'Ẽ', 'Ì', 'Í', 'Î', 'Ï', 'Ð', 'Ñ', 'Ò', 'Ó', 'Ô', 'Õ', 'Ö', 'Ø', 'Ù', 'Ú', 'Û', 'Ü', 'Ý','Ç', 'ß', 'à', 'á', 'â', 'ã', 'ä', 'å', 'æ', 'ç', 'è', 'é', 'ê', 'ë', 'ẽ', 'ì', 'í', 'î', 'ï', 'ñ', 'ò', 'ó', 'ô', 'õ', 'ö', 'ø', 'ù', 'ú', 'û', 'ü', 'ý', 'ÿ', 'Ā', 'ā', 'Ă', 'ă', 'Ą', 'ą', 'Ć', 'ć', 'Ĉ', 'ĉ', 'Ċ', 'ċ', 'Č', 'č', 'Ď', 'ď', 'Đ', 'đ', 'Ē', 'ē', 'Ĕ', 'ĕ', 'Ė', 'ė', 'Ę', 'ę', 'Ě', 'ě', 'Ĝ', 'ĝ', 'Ğ', 'ğ', 'Ġ', 'ġ', 'Ģ', 'ģ', 'Ĥ', 'ĥ', 'Ħ', 'ħ', 'Ĩ', 'ĩ', 'Ī', 'ī', 'Ĭ', 'ĭ', 'Į', 'į', 'İ', 'ı', 'IJ', 'ij', 'Ĵ', 'ĵ', 'Ķ', 'ķ', 'Ĺ', 'ĺ', 'Ļ', 'ļ', 'Ľ', 'ľ', 'Ŀ', 'ŀ', 'Ł', 'ł', 'Ń', 'ń', 'Ņ', 'ņ', 'Ň', 'ň', 'ʼn', 'Ō', 'ō', 'Ŏ', 'ŏ', 'Ő', 'ő', 'Œ', 'œ', 'Ŕ', 'ŕ', 'Ŗ', 'ŗ', 'Ř', 'ř', 'Ś', 'ś', 'Ŝ', 'ŝ', 'Ş', 'ş', 'Š', 'š', 'Ţ', 'ţ', 'Ť', 'ť', 'Ŧ', 'ŧ', 'Ũ', 'ũ', 'Ū', 'ū', 'Ŭ', 'ŭ', 'Ů', 'ů', 'Ű', 'ű', 'Ų', 'ų', 'Ŵ', 'ŵ', 'Ŷ', 'ŷ', 'Ÿ', 'Ź', 'ź', 'Ż', 'ż', 'Ž', 'ž', 'ſ', 'ƒ', 'Ơ', 'ơ', 'Ư', 'ư', 'Ǎ', 'ǎ', 'Ǐ', 'ǐ', 'Ǒ', 'ǒ', 'Ǔ', 'ǔ', 'Ǖ', 'ǖ', 'Ǘ', 'ǘ', 'Ǚ', 'ǚ', 'Ǜ', 'ǜ', 'Ǻ', 'ǻ', 'Ǽ', 'ǽ', 'Ǿ', 'ǿ', 'Ά', 'ά', 'Έ', 'έ', 'Ό', 'ό', 'Ώ', 'ώ', 'Ί', 'ί', 'ϊ', 'ΐ', 'Ύ', 'ύ', 'ϋ', 'ΰ', 'Ή', 'ή', '…', '.', '\'', ',');
  $b = array('A', 'A', 'A', 'A', 'A', 'A', 'AE', 'C', 'E', 'E', 'E', 'E', 'E', 'I', 'I', 'I', 'I', 'D', 'N', 'O', 'O', 'O', 'O', 'O', 'O', 'U', 'U', 'U', 'U', 'Y','C', 's', 'a', 'a', 'a', 'a', 'a', 'a', 'ae', 'c', 'e', 'e', 'e', 'e', 'e', 'i', 'i', 'i', 'i', 'n', 'o', 'o', 'o', 'o', 'o', 'o', 'u', 'u', 'u', 'u', 'y', 'y', 'A', 'a', 'A', 'a', 'A', 'a', 'C', 'c', 'C', 'c', 'C', 'c', 'C', 'c', 'D', 'd', 'D', 'd', 'E', 'e', 'E', 'e', 'E', 'e', 'E', 'e', 'E', 'e', 'G', 'g', 'G', 'g', 'G', 'g', 'G', 'g', 'H', 'h', 'H', 'h', 'I', 'i', 'I', 'i', 'I', 'i', 'I', 'i', 'I', 'i', 'IJ', 'ij', 'J', 'j', 'K', 'k', 'L', 'l', 'L', 'l', 'L', 'l', 'L', 'l', 'l', 'l', 'N', 'n', 'N', 'n', 'N', 'n', 'n', 'O', 'o', 'O', 'o', 'O', 'o', 'OE', 'oe', 'R', 'r', 'R', 'r', 'R', 'r', 'S', 's', 'S', 's', 'S', 's', 'S', 's', 'T', 't', 'T', 't', 'T', 't', 'U', 'u', 'U', 'u', 'U', 'u', 'U', 'u', 'U', 'u', 'U', 'u', 'W', 'w', 'Y', 'y', 'Y', 'Z', 'z', 'Z', 'z', 'Z', 'z', 's', 'f', 'O', 'o', 'U', 'u', 'A', 'a', 'I', 'i', 'O', 'o', 'U', 'u', 'U', 'u', 'U', 'u', 'U', 'u', 'U', 'u', 'A', 'a', 'AE', 'ae', 'O', 'o', 'Α', 'α', 'Ε', 'ε', 'Ο', 'ο', 'Ω', 'ω', 'Ι', 'ι', 'ι', 'ι', 'Υ', 'υ', 'υ', 'υ', 'Η', 'η', '', '', '', '');
  return str_replace($a, $b, $str);
}

if(isset($_POST['heya2'])) {
  echo removeAccents($_POST['heya']);
  echo '<br>';
  echo strtr($_POST['heya'],'àáâãäçèéêëìíîïñòóôõöùúûüýÿÀÁÂÃÄÇÈÉÊËÌÍÎÏÑÒÓÔÕÖÙÚÛÜÝ','aaaaaceeeeiiiinooooouuuuyyAAAAACEEEEIIIINOOOOOUUUUY');
}

HTML:

<form action="" method="POST">
  <input name="heya" type="text">
  <input name="heya2" type="submit">
</form>

In the above example, the output of the two "Indian" functions is exactly the same, "Indian".

In the function to modify the database happens the same:

$dataBase = new DB($db);

$projects = $dataBase->fetchAllProjectsByDisplayOrder();
foreach($projects as $p) {
  $shortName = strtolower(str_replace('_','-', removeAccents($p->short_name)));
  //$dataBase->updateShortNameMigra2($p->id, $shortName);
  echo $shortName. '<br>';
}

Below is an image of what happens with this code. Some way to resolve this, remove the accents?

printscreen

  • Why remove accents? It is to generate SEO links or other specific thing?

  • Exact is for the URL to be cleaner. Instead of ...project.php?name=Ao_Ritmo_Da_ProduÇÃo for example

  • i tested your script the removeAccents() function worked here. Only strtr that gave problem.

  • I really don’t notice. The function results if we clear the string directly, echo strtolower(str_replace('_','-', removeAccents('Inês')));, but if we write it ...strtolower(str_replace('_','-', removeAccents($_POST['heya']))); it’s no longer possible. I saw now that when grabbing the database also happens the same, if we write "Ao_ritmo_da_production" directly works, but not as variable

  • With iconv it’s the same if $output = 'Ao_Ritmo_Da_ProduÇÃo';&#xA;$foo = iconv('UTF-8','ASCII//TRANSLIT',$output);, results, but if the variable does not come directly from the script no longer... ...iconv('UTF-8','ASCII//TRANSLIT',$p->short_name); // neste caso $p->short_name = Ao_Ritmo_Da_ProduÇÃo it is no longer possible

  • make sure on the charset settings of both the data in the database and the php scripts.

  • It’s all right (UTF-8)

  • Guys I found out what was going on, it’s down. Thank you all

Show 3 more comments

1 answer

0


I did it. After thinking and thinking I discovered that as I sent the data to the database as HTML entities with the function htmlentities() by security (eg: htmlentities(só) = s&oacute;) when I fetch the data from the database they have no accents, ie the function removeAccents() it did nothing because no word has accents, they only came to exist in the browser where this function was no longer executed, I have to do the Decode (html_entity_decode()) on the server first and take out the accents later:

foreach($projects as $p) {
  echo removeAccents(html_entity_decode($p->short_name)). '<br>';
}

Browser other questions tagged

You are not signed in. Login or sign up in order to post.