UTF-32/UCS-4
Sign in to saveAlso known as 32-bit Unicode Transformation Format, Unicode Transformation Format – 32-bit, Unicode Transformation Format - 32-bit, UTF32, UTF_32, UCS-4, UCS4
UTF-32 (32-bit Unicode Transformation Format), sometimes called UCS-4, is a fixed-length encoding used to encode Unicode code points that uses exactly 32 bits (four bytes) per code point (but a number of leading bits must be zero as there are far fewer than 232 Unicode code points, needing actually only 21 bits). In contrast, all other Unicode transformation formats are variable-length encodings. Each 32-bit value in UTF-32 represents one Unicode code point and is exactly equal to that code point's numerical value.
Described at

Note: This feedback goes to your product's documentation team and does not include a response. Issues that require a response should go through IBM support. Note: This feedback goes to your product's documentation team and does not include a response. Issues that require a response should go through IBM support. Note: This feedback goes to your product's documentation team and does not include a response. Issues that require a response should go through IBM support. There has been an error sending your feedback to the team. Your comment was saved locally, if not in an incognito browser, and will be available when attempting to submit feedback again. The IBM® i operating system does not support UTF-32 encoding with a CCSID value. Unicode was originally designed as a pure 16-bit encoding, aimed at representing all modern scripts. Over time, and especially after the addition of over 14 500 composite characters for compatibility with established sets, it became clear that 16 bits were not sufficient for many users. Out of this arose UTF-32. Arrow rightDifferent encodings of Unicode is the algorithmic mapping from every Unicode value to a unique byte sequence.") While IBM values the use of inclusive language, terms that are outside of IBM's direct influence, for the sake of maintaining user understanding, are sometimes required. As other industry leaders join IBM in embracing the use of inclusive language, IBM will continue to update the documentation to reflect those changes.
Excerpt from a page describing this subject · 12,470 chars · not written by Vinony
Wikidata facts
Show 4 more facts
- Stack Exchange tag
- stackoverflow.com/tags/utf-32
- different from
- Unicode
- described at URL
- www.ibm.com/docs/en/i/7.1?topic=unicode-utf-32
- data size
- 32
Sources (1)
via Wikidata · CC0
Article · Português
UTF-32 ou UCS-4 são nomes alternativos para o método de codificação de caracters, usando a quantidade fixa de exatamente 32 bits para cada caractere Unicode. Ele pode ser considerado como a forma de codificação mais simples, como todos os outros Unicode Transformation Formats (em português: Formato de Transformação Unicode) possui codificação de comprimento variável para vários code points. No entanto, o UTF-32 usa 4 bytes para cada caractere, que é considerado ineficiente. Especificamente, caracteres que não pertencem ao (PBM) são tão raros em quase todos os textos que eles podem ser considerados como pouco importantes para discussões importantes. Isto significa que UTF-32 é geralmente pelo menos o dobro ou quatro vezes maior que o tamanho normal das outras codificações. Também, enquanto um número fixo de bytes por ponto de código pareça ser conveniente de primeiro, não é. Torna o truncamento levemente mais fácil, mas não tão significativo de UTF-8 e UTF-16. Não faz o cálculo de largura de uma string exibida mais fácil, exceto em casos muito limitados; mesmo com uma fonte de "tamanho fixo" pode haver mais que um ponto de código por posição de caractere (marcas combinadas) (por exemplo ideógrafos CJK). Combinando marcas também quer dizer que os editores não podem tratar um ponto de código como se fosse uma unidade para edição. Por estas razões o UTF-32 é pouco utilizado na prática, com UTF-8 e UTF-16 sendo o método comum de codificar texto Unicode.
Abstract from DBpedia / Wikipedia · CC BY-SA